Performs hour-scale video understanding and fine-grained temporal localization while exposing agent-style multimodal tool/code/search abilities. Built on a sparse-attention long-context architecture (DSA) and a specialized inference stack—best used in GPU-backed research or production evaluation.
Provides the renderer weights and inference code for Bernini’s video renderer, enabling text→video, image→video and video editing inference. Offers a ready diffusers-format bundle or safetensors checkpoints under Apache‑2.0; intended for multi‑GPU/Hopper inference and reproducible research.
Generates minute-level, multi-shot synchronized audio+video from a single text prompt, using a paired cross-modal memory to preserve character appearance and voice across shots. Uses DMD-distilled few-step inference for ~7.5× speedup; requires high-GPU memory and is released under the LTX-2 community license.
End-to-end pose-driven image-to-video model that animates a reference character from a driving video, supporting cross-identity replacement and multi-character scenarios without intermediate pose representations; performs best at 704p and ships as a diffusers-compatible checkpoint.
Converts low‑poly 3D viewport or game/CG renders into photorealistic cinematic video while preserving the input's composition, camera motion and layout; offers Light and Strong LoRA variants to trade fidelity for aggressive photorealism.
Generates short videos that preserve a reference person's identity from a single reference image as a LoRA adapter for LTX-2. Uses overlap reference conditioning with TASS‑RoPE source-phase tagging and an ArcFace identity loss; runs in ComfyUI via BFS Nodes and supports a 4‑panel character‑sheet mode for clothing/body consistency.
Converts an academic paper into reusable extracted assets and then produces editable poster, synchronized talk video, and bilingual blog via modular generator skills. Key differentiator: a single Paper2Assets extractor shared by three editable generators plus an interactive Paper2Reel viewer that links slides, video, captions and blog while preserving factual consistency and round-tripable PPT/DOCX output.
Generates image-to-video world-model outputs using a distilled 14B causal model optimized for chunked, KV-cached inference across long-horizon interactive scenes; offers a real-time 'causal-fast' variant capable of driving near‑real‑time video streams and an agentic harness for action-driven scene synthesis (CC BY‑NC‑SA).
Generates videos from text and image+text prompts using a 30B Mixture-of-Experts model tuned for embodied intelligence; includes a refiner and structured prompt rewriter, and supports diffusers/SGLang runtimes with multi-GPU inference.
Generates minute-scale, temporally coherent dance videos from full music tracks using a hierarchical two-stage approach: global keyframe planning plus local temporal refinement; suitable when long-range musical structure and rhythmic continuity matter.
Generates a new camera viewpoint from a reference video: an IC‑LoRA adapter for LTX‑Video 2.3 that re‑renders the same scene from a requested discrete camera angle while preserving subject and content. Trained on synthetic multi‑view data, proof‑of‑concept with limited viewpoint range and best for small, chained angle shifts.
Generates synchronized audiovisual output from text, image, or audio prompts — a diffusion-based multimodal model with componentized weights (video/audio VAEs, multilingual text encoder, distilled transformer) and ready integration with HuggingFace pipelines and ComfyUI.