An agentic framework that analyzes, plans, and executes multi-step video understanding and editing workflows using multimodal LLM-driven agents—features intent decomposition, graph-based workflow orchestration, and automated shot planning for long-form video tasks.
Framework for building multi-modal AI agents that watch, listen, and reason over live video, pairing vision models (YOLO, Roboflow, Moondream) with LLMs like Gemini and OpenAI. Agents join calls in ~500ms and keep audio/video latency under 30ms.
Generates explorable, 3D-consistent virtual worlds from a single image or short video. Includes official implementations of Lyra‑1 (feed‑forward 3D/4D scene generation via video-diffusion self-distillation) and Lyra‑2 (long-horizon, explorable generative 3D worlds). Best for research and creative prototyping; requires substantial GPU compute.
Provides an NVFP4‑optimized training and inference infrastructure for long-form video diffusion models — supports multi-shot AR training, KV-cache and NVFP4 quantized inference, sequence-parallelism and async decoding for higher FPS and longer outputs.
Automatically generates complete short-form videos from a single topic: drafts script with an LLM, produces AI images/video, synthesizes multilingual TTS (including voice cloning), adds background music, and composes the final video. Supports local ComfyUI/RunningHub or direct model APIs and customizable templates.
Runs text-to-video, image-to-video, text-to-image, and image editing inference with acceleration, offloading, quantization, and distributed execution for large visual generation models.
Generates real-time, infinite-length portrait video from one reference image on a 12GB GPU. Combines implicit facial signals and 3D keypoints with step-distilled diffusion and autoregressive micro-chunk streaming for low-latency live use.
Provides a DiT-based audio–video foundation model plus an official Python inference and LoRA trainer. Ships multiple production-ready pipelines (text/image/audio→video), checkpoints, and performance optimizations (FP8, distilled pipelines) for high-fidelity synchronized audio–video generation.
Browser-based, client-side video editor for multi-track editing, GPU-accelerated preview and local exports without uploading files; leverages WebCodecs/WebGPU and includes an AI upscaling option.
Turns natural-language directions into end-to-end video editing workflows: LLM-powered planning, media search/organization, ASR rough-cut, and reusable Style Skills for consistent storytelling. Integrates agent Skills (OpenClaw/Claude Code) and optional AIGC transitions.
Unmixes green‑screen pixels with a neural model to recover straight (unmultiplied) foreground color and a clean linear alpha for every pixel, preserving hair, motion blur and translucency. Produces VFX‑standard EXR outputs, supports optional AlphaHint generators (GVM/VideoMaMa) and Docker/consumer‑GPU optimizations.
Converts scene intent into production-ready Seedance 2.0 prompts, reference-role mappings, and IP-safe rewrites for multimodal (text/image/audio/video) video generation. Ships as a modular agent-skill OS with multilingual examples, troubleshooting tools, and pro filmmaker handoff artifacts.