Generates temporally grounded captions for dense multi-event videos by restructuring autoregressive token dependencies to enable lossless parallel decoding; introduces a latent global planning module and event-factorized parallel decoding to improve grounding accuracy and achieve large decoding speedups.
Generates real-time, infinite-length interactive videos of voice-controllable digital characters — 540p at up to 42 FPS on consumer GPUs. Uses TurboDiffusion and TurboServe to maintain temporal coherence without blur or drift, and accepts custom person, anime, or pet images plus selectable voice tones.
Generates short videos that preserve a reference person's identity from a single reference image as a LoRA adapter for LTX-2. Uses overlap reference conditioning with TASS‑RoPE source-phase tagging and an ArcFace identity loss; runs in ComfyUI via BFS Nodes and supports a 4‑panel character‑sheet mode for clothing/body consistency.
Converts an academic paper into reusable extracted assets and then produces editable poster, synchronized talk video, and bilingual blog via modular generator skills. Key differentiator: a single Paper2Assets extractor shared by three editable generators plus an interactive Paper2Reel viewer that links slides, video, captions and blog while preserving factual consistency and round-tripable PPT/DOCX output.
Provides re-annotated academic video instruction data for captioning, video QA, and fine-grained motion understanding; rewrites short answers and concise captions into evidence-grounded, instruction-following responses and supplies JSONL annotation files (original videos not included).
Provides a reflexive agentic framework for long-horizon video understanding that replaces costly iterative reasoning with dual contextual states: a consolidated global multimodal script and parametric latent states for fast retrieval and response, improving speed and memory efficiency.
Provides structured egocentric manipulation signals from smartphone videos: MANO 3D hand reconstructions, metric camera trajectories, and fine-grained atomic action segments (full release ≈2,000 hours planned). Supplies aligned hands.npz, camera_traj.npz, undistorted intrinsics and segment annotations for embodied-learning pipelines.
Autoregressively synthesizes long-horizon, playable video worlds conditioned on current state and user actions for real-time interaction. Ships as an open-source, full-stack framework covering data preparation, model architectures, training, inference acceleration, and deployment for interactive generative worlds.
Generates image-to-video world-model outputs using a distilled 14B causal model optimized for chunked, KV-cached inference across long-horizon interactive scenes; offers a real-time 'causal-fast' variant capable of driving near‑real‑time video streams and an agentic harness for action-driven scene synthesis (CC BY‑NC‑SA).
Pretrains a DiT-based Mixture-of-Experts video foundation model for embodied intelligence by augmenting internet videos with robot-centric footage and using a multi-dimensional reward system to prioritize physical realism and task completion while scaling MoE for better capacity vs. inference trade-offs.
Generates videos from text and image+text prompts using a 30B Mixture-of-Experts model tuned for embodied intelligence; includes a refiner and structured prompt rewriter, and supports diffusers/SGLang runtimes with multi-GPU inference.
Creates an open-ended interactive world simulator with an unbounded interaction horizon via causal pretraining, a distilled real-time runtime that drives 720p@60fps, a wider action/event repertoire, and a pilot–director agent split for behavior planning and environment synthesis.