Generates real-time, infinite-length portrait video from one reference image on a 12GB GPU. Combines implicit facial signals and 3D keypoints with step-distilled diffusion and autoregressive micro-chunk streaming for low-latency live use.
Generates summaries from URLs, YouTube videos, podcasts, PDFs, and local audio or video files. Backend-agnostic by design: the same pipeline drives local coding CLIs (Claude, Codex, Gemini) or hosted API providers (OpenAI, Google, xAI).
Large-scale, real-world dual-arm video corpus for embodied robotics and reinforcement-learning research — over 1TB of multimodal recordings on Hugging Face, intended for training and evaluating agents in real manipulation scenarios; CC BY‑SA 4.0.
Provides a DiT-based audio–video foundation model plus an official Python inference and LoRA trainer. Ships multiple production-ready pipelines (text/image/audio→video), checkpoints, and performance optimizations (FP8, distilled pipelines) for high-fidelity synchronized audio–video generation.
Browser-based, client-side video editor for multi-track editing, GPU-accelerated preview and local exports without uploading files; leverages WebCodecs/WebGPU and includes an AI upscaling option.
Generates complete Godot 4 projects from a natural-language game description: it designs the architecture, generates assets, writes C# code, runs the project, captures screenshots for visual QA, and iterates until a runnable game repo is produced. Requires API keys and Godot .NET.
Turns natural-language directions into end-to-end video editing workflows: LLM-powered planning, media search/organization, ASR rough-cut, and reusable Style Skills for consistent storytelling. Integrates agent Skills (OpenClaw/Claude Code) and optional AIGC transitions.
Lets any LLM operate a ComfyUI instance: generate and iterate images/video/audio, manage models and custom nodes, and edit the live graph in natural language. Local-first control plane with a sidebar agent, multi-provider LLM support, and installer packs for ready workflows.
Unmixes green‑screen pixels with a neural model to recover straight (unmultiplied) foreground color and a clean linear alpha for every pixel, preserving hair, motion blur and translucency. Produces VFX‑standard EXR outputs, supports optional AlphaHint generators (GVM/VideoMaMa) and Docker/consumer‑GPU optimizations.
Converts scene intent into production-ready Seedance 2.0 prompts, reference-role mappings, and IP-safe rewrites for multimodal (text/image/audio/video) video generation. Ships as a modular agent-skill OS with multilingual examples, troubleshooting tools, and pro filmmaker handoff artifacts.
ComfyUI workflows that run LTX‑2.3 split models to produce text→video, image→video and audio→video pipelines. Uses extracted/split safetensor or GGUF files so models load more modularly; requires up‑to‑date ComfyUI, KJNodes and ComfyUI‑GGUF.
Author HTML-based video compositions and render deterministic, frame-accurate MP4s with agent-friendly tooling — preview in the browser, drive generation via AI agent skills, and use adapter runtimes (GSAP, Lottie, Three.js).