Creates an open-ended interactive world simulator with an unbounded interaction horizon via causal pretraining, a distilled real-time runtime that drives 720p@60fps, a wider action/event repertoire, and a pilot–director agent split for behavior planning and environment synthesis.
Recovers and predicts RGB video from sparse event-camera streams by fine-tuning pre-trained video diffusion priors; jointly addresses reconstruction, long-horizon prediction, and bidirectional frame interpolation with mechanisms to reduce temporal drift and enforce interpolation consistency.
Uses large-scale text-to-video generative pretraining to create GenCeption, a feed-forward perception model that performs diverse vision tasks from text instructions—depth, surface normals, camera pose, referring segmentation, and 3D keypoints—often matching or surpassing specialized models while requiring far less task-specific data.
Reconstructs 4D dynamic human scenes from sparse, low-overlap multi-camera captures by decoupling background synthesis and human modeling. Synthesizes hundreds of camera-controlled background views with a video diffusion model, initializes deformable Gaussian humans via cross-view identity and triangulated keypoints, then applies motion-adaptive recursive enhancement to reduce artifacts.
Generates minute-scale, temporally coherent dance videos from full music tracks using a hierarchical two-stage approach: global keyframe planning plus local temporal refinement; suitable when long-range musical structure and rhythmic continuity matter.
Generates a new camera viewpoint from a reference video: an IC‑LoRA adapter for LTX‑Video 2.3 that re‑renders the same scene from a requested discrete camera angle while preserving subject and content. Trained on synthetic multi‑view data, proof‑of‑concept with limited viewpoint range and best for small, chained angle shifts.
Timestamp-aware realtime video→text model that processes incoming frames continuously, answers questions mid-stream or emits silence when evidence is insufficient, and can revise earlier outputs as new frames arrive. Built for timestamped multimodal interaction with a 256K context and an 11B-parameter backbone.
Comprehensive benchmark and automated evaluation framework for keyframe-conditioned video generation—decomposes keyframe execution into six metrics and assesses overall video quality with evidence-grounded MLLM judgments and specialized perception models.
Enables efficient, generalist video understanding by combining an Inflated 3D Vision Transformer and adaptive frame-resolution streaming with a scalable video data synthesis pipeline; ships as a fully open 4B-parameter MLLM that improves general, long-form, and streaming benchmarks.
Evaluates whether video models reason according to physical laws by treating generated videos as visible reasoning traces and using a three-stage Perception–Formulation–Deduction protocol. Includes Orchard (400 mechanics videos), chain-of-frames prompting on annotated first frames, and a hybrid MLLM-plus-objective scoring suite for stage-resolved diagnostics.
Predicts variable-cardinality sets of evidence intervals in videos to temporally ground queries using multimodal large language models. Combines caption-derived multi-span supervision, a temporal Wasserstein matching-free reward, and temporal IoU, yielding strong mIoU gains across multiple benchmarks.
Provides 2,000 hours of synchronized, high‑fidelity robot‑free bimanual manipulation demonstrations with multi‑view video, calibrated end‑effector trajectories, gripper states, and language annotations. Curated from a 20,000+ hour corpus; features 6 camera views, ~3 mm pose accuracy, <40 µs cross‑sensor sync, and LeRobot v3‑style Parquet+MP4 export under CC BY 4.0.