Enables real-time streaming video-to-video editing (1280×704 @24 FPS) on a single RTX 5090 GPU. Uses a Hybrid Diffusion Transformer for balanced local/global modeling, Cycle‑Reverse Regularization for temporal consistency, and system-level mixed-precision and fused kernels to maximize throughput.
Generates synchronized, streaming spatial audio from panoramic video and text prompts using a causal autoregressive diffusion transformer. Combines Spatial Video-Audio Contrastive (SVAC) alignment and online direct preference optimization (ODPO) to improve spatial perception, plus an automated annotation pipeline and public demos.
Provides the renderer weights and inference code for Bernini’s video renderer, enabling text→video, image→video and video editing inference. Offers a ready diffusers-format bundle or safetensors checkpoints under Apache‑2.0; intended for multi‑GPU/Hopper inference and reproducible research.
Omnimodal world model that jointly processes and generates text, images, video, audio, and action trajectories for physical AI. Uses a mixture-of-transformers to combine autoregressive reasoning and diffusion-based multimodal generation; released open-source with checkpoints, datasets and benchmarks for robotics and simulation.
Evaluates multimodal LLMs on streaming egocentric video for spatial intelligence using 1,680 human-annotated questions across 348 videos; organizes tasks into four hierarchical levels (perception → tracking → simulation → allocentric mapping) and highlights allocentric mapping as the main bottleneck.
Generates minute-level, multi-shot synchronized audio+video from a single text prompt, using a paired cross-modal memory to preserve character appearance and voice across shots. Uses DMD-distilled few-step inference for ~7.5× speedup; requires high-GPU memory and is released under the LTX-2 community license.
Native multimodal model for image/text/video→text tasks with million‑token context support. Uses a sparse-attention operator to cut long‑context compute and latency, and targets agentic, coding, and long-horizon conversational workloads.
Large-scale training corpus for knowledge- and reasoning-intensive video understanding: 315K video reasoning examples over 145K CC-licensed expert-domain videos, with human-in-the-loop chain-of-thought rationales to strengthen post-training for video reasoning. ([arxiv.org](https://arxiv.org/abs/2606.05259))
Provides 1,036,431 identity–text–video triplets with per-video JSON annotations and reference keyframes to train and evaluate identity-preserving customized video generation models. Data is drawn from ~320K Pexels HD videos; videos must be downloaded separately per Pexels' terms.
Stores a persistent 3D scene cache directly in a diffusion model's latent space to produce temporally and spatially consistent videos. Constructs memory via depth-guided back-projection and queries it with direct latent-space warping — achieving large speed and memory gains versus pixel-space 3D baselines.
End-to-end framework for controlled character animation that transfers motion from driving videos to reference characters without intermediate pose or background representations. Introduces the MotionPair‑60K end-to-end motion-transfer dataset, in‑context mask conditioning and mode‑specific RoPE for task unification, plus Bias‑Aware DPO to mitigate synthetic-detail errors.
Provides 500+ hours of human whole-body teleoperation demonstrations for humanoid robot learning in real homes, with synchronized video, joint states, action traces and language annotations. Includes 23K+ episodes, fine-grained subtask labels, and raw ROS/MCAP plus compressed LeRobot formats.