Serves interactive, long-lived streaming video-generation sessions by jointly scheduling session placement and GPU autoscaling to meet tight per-chunk latency. Combines migration-aware placement, load-driven autoscaling, coalesced chunk processing, GPU–CPU offloading and NCCL GPU–GPU migration; reports ~37% reductions in worst-case per-chunk latency and GPU operating cost.
Provides a small, manually annotated benchmark for evaluating vision–language models that convert robot and egocentric manipulation videos into timestamped subtask segments and concise action labels. Contains 100 episodes, 743 gold segments, and MP4 bytes embedded per row.
Provides multiview synthetic RGB video clips with per-frame depth, instance masks, dense long-range 3D point tracks, camera poses, and SMPL‑X human pose/shape labels for 4D reconstruction, tracking, and geometry-aware novel-view synthesis. Includes ~4.7K clips (1.4M frames) and is licensed for AI training.
Provides ~2 million instruction-aligned video-edit pairs for training and evaluating instruction-based video editing and generation models. Covers multi-task and structural edits (e.g., camera/subject movement), produced via a synthesis pipeline with progressive filtering; licensed CC BY-NC-4.0.
Multimodal video dataset for text-to-video and video-to-video research: about 2 million short English videos and extracted frames for instruction-based video editing and generation. Hosted on Hugging Face and licensed CC BY‑NC 4.0 (non-commercial).
Provides 600,000+ first-person player-round videos (10,000+ hours) with per-frame keyboard, mouse-delta, and 3D trajectory annotations in WebDataset shards—built for training world models, action-conditioned video, and imitation-learning workflows (non-commercial license).
Provides synchronized four-perspective Rocket League match recordings with per-frame H.264 video, player action streams, event logs, and privileged physics state — released as WebDataset shards in a ~4,000-hour slice (1,000 match-hours × 4 perspectives). Includes 720p@20fps video, multi-hot keyboard actions, and CC BY-NC-SA-4.0 license.
Provides ~494.7 hours of trimmed native PC/console gameplay screen recordings organized by game, with per-session clips plus input and per-frame event annotations. Each workflow includes clip.mp4, events.json, frame_events.json, and metadata — suitable for training vision-action, behavior-cloning, and gameplay understanding models.
Provides anonymized multi-domain user behavior sequences and content metadata (short video, ads, e-commerce, live) for cross-domain recommendation, semantic-ID mapping, and content-understanding tasks. Key tables include per-user multi-domain behavior (~500k rows), pid→three-segment semantic IDs, captions, and level-3 tags; all item IDs are hashed for privacy.
Diffusion-based generative model for scene and video synthesis, providing full Diffusers checkpoints and scene LoRA for fast adaptation. Includes Stage‑1 nano (1.3B) and pro (5B) variants and modular transformer/VAE components.
Provides a systematic benchmark and design roadmap for video-based world models to evaluate robot policies, introducing WMBench and GigaWorld-1 optimized for long-horizon, action-faithful rollouts. Offers controlled comparisons across model families, action encodings, and 324k+ simulated vs real rollouts, with code, models, and datasets released for reproducible evaluation.
Generates temporally grounded captions for dense multi-event videos by restructuring autoregressive token dependencies to enable lossless parallel decoding; introduces a latent global planning module and event-factorized parallel decoding to improve grounding accuracy and achieve large decoding speedups.