Generates real-time, infinite-length interactive videos of voice-controllable digital characters — 540p at up to 42 FPS on consumer GPUs. Uses TurboDiffusion and TurboServe to maintain temporal coherence without blur or drift, and accepts custom person, anime, or pet images plus selectable voice tones.
Converts an academic paper into reusable extracted assets and then produces editable poster, synchronized talk video, and bilingual blog via modular generator skills. Key differentiator: a single Paper2Assets extractor shared by three editable generators plus an interactive Paper2Reel viewer that links slides, video, captions and blog while preserving factual consistency and round-tripable PPT/DOCX output.
Provides a reflexive agentic framework for long-horizon video understanding that replaces costly iterative reasoning with dual contextual states: a consolidated global multimodal script and parametric latent states for fast retrieval and response, improving speed and memory efficiency.
Autoregressively synthesizes long-horizon, playable video worlds conditioned on current state and user actions for real-time interaction. Ships as an open-source, full-stack framework covering data preparation, model architectures, training, inference acceleration, and deployment for interactive generative worlds.
Pretrains a DiT-based Mixture-of-Experts video foundation model for embodied intelligence by augmenting internet videos with robot-centric footage and using a multi-dimensional reward system to prioritize physical realism and task completion while scaling MoE for better capacity vs. inference trade-offs.
Recovers and predicts RGB video from sparse event-camera streams by fine-tuning pre-trained video diffusion priors; jointly addresses reconstruction, long-horizon prediction, and bidirectional frame interpolation with mechanisms to reduce temporal drift and enforce interpolation consistency.
Uses large-scale text-to-video generative pretraining to create GenCeption, a feed-forward perception model that performs diverse vision tasks from text instructions—depth, surface normals, camera pose, referring segmentation, and 3D keypoints—often matching or surpassing specialized models while requiring far less task-specific data.
Comprehensive benchmark and automated evaluation framework for keyframe-conditioned video generation—decomposes keyframe execution into six metrics and assesses overall video quality with evidence-grounded MLLM judgments and specialized perception models.
Enables efficient, generalist video understanding by combining an Inflated 3D Vision Transformer and adaptive frame-resolution streaming with a scalable video data synthesis pipeline; ships as a fully open 4B-parameter MLLM that improves general, long-form, and streaming benchmarks.
Evaluates whether video models reason according to physical laws by treating generated videos as visible reasoning traces and using a three-stage Perception–Formulation–Deduction protocol. Includes Orchard (400 mechanics videos), chain-of-frames prompting on annotated first frames, and a hybrid MLLM-plus-objective scoring suite for stage-resolved diagnostics.
Predicts variable-cardinality sets of evidence intervals in videos to temporally ground queries using multimodal large language models. Combines caption-derived multi-span supervision, a temporal Wasserstein matching-free reward, and temporal IoU, yielding strong mIoU gains across multiple benchmarks.
Personalizes subject-driven videos to preserve human identity and accurate human–object interactions by integrating multimodal references and MLLM-derived semantics. Introduces global multimodal guidance in self-attention and modality-reference embeddings to align MLLM features with VAE tokens, supporting both inter- and intra-subject inputs (e.g., OCR, multi-view).