Provides 1.3 billion platform-specific video URLs extracted from CommonCrawl along with crawl metadata (no media included), serving as the source corpus for the LAION-BVD multimodal video dataset; distributed on Hugging Face in Parquet format.
Generates uncensored videos from text and images using an LTX 2.3–based diffusion model with native t2v and i2v support; ships with a prompt enhancer and developer-focused gguf/bf16 dev releases for local experimentation.
Large-scale synthetic video dataset of physically simulated multi-object interaction scenes for training and evaluating models on physical reasoning, depth and optical-flow estimation, instance segmentation, and physics-grounded captioning. Provides RGB + lossless depth, per-frame instance masks, per-object physics annotations (NPZ), VLM-grounded captions, and USD scene files — useful for world-model and simulation-to-real work; commercial use permitted.
Provides 55 million scene-level video clips (each with captions, language labels, and timestamps) extracted from an 80M-video, 10-million-hour raw pool to support multimodal pre-training across video, audio, and frames. Access is gated for academic/non-commercial research.
Open egocentric multimodal dataset for embodied AI and robot learning captured on commodity iPhone Pro: ~200 hours and ~10M RGB frames with LiDAR depth, ARKit 6‑DoF poses, IMU, two‑hand MANO mocap, room meshes, and hierarchical action captions.
Provides tick-aligned Counter-Strike 2 player POV video clips with per-tick inputs and world-state sidecars — near-lossless 1280×720@32fps video, per-player stereo audio, and parquet indexes for event/kill/round filtering; suited for RL, video classification and clip mining.
Converts video inputs into text outputs — supports captioning, temporal grounding, and video-text-to-text queries using a Qwen-3.5-2B finetuned multimodal backbone. Suited for prototyping video understanding and caption-generation pipelines.
Delivers image and video generation, editing, and understanding inside a single 3B-parameter multimodal model trained from scratch with a multi-task recipe. Notable for strong unified benchmarks at 3B scale; inference requires large GPU memory (≈40GB+ VRAM).
Generates minute-scale, 720p videos from a single image using a 2.6B image-to-video diffusion transformer with precise 6‑DoF camera control and an optional LTX‑2 refiner; designed for long-context, memory-efficient modeling but requires large refiner checkpoints (~41 GB).
Generates temporally coherent MP4 videos from a single input image plus text instructions, with configurable resolution, frame count, and optional AAC audio. Optimized for NVIDIA GPU stacks and integrates with vLLM‑Omni and Hugging Face Diffusers for production inference and research workflows.
Generates audio-driven avatar videos from text, images, or audio inputs with production-grade stability (accurate lip sync, identity consistency) and an 8-step distillation inference mode for faster serving; suitable for broadcasting, virtual hosts, animation, and multi-person scenarios.
Terminal-native AI coding agent that reads and edits code, runs shell commands, searches files, fetches web pages, and determines next steps from interactive feedback. Delivered as a single-binary TUI with video input, subagents, a plugin marketplace, and IDE (ACP) integration.