ComfyUI workflows that run LTX‑2.3 split models to produce text→video, image→video and audio→video pipelines. Uses extracted/split safetensor or GGUF files so models load more modularly; requires up‑to‑date ComfyUI, KJNodes and ComfyUI‑GGUF.
Author HTML-based video compositions and render deterministic, frame-accurate MP4s with agent-friendly tooling — preview in the browser, drive generation via AI agent skills, and use adapter runtimes (GSAP, Lottie, Three.js).
Orchestrates end-to-end video production with agentic pipelines that research, script, generate assets, edit, and render finished videos. Distinguishes itself by supporting true real-footage retrieval (Archive.org, NASA, Wikimedia), Remotion/HyperFrames composition, and usable zero-key workflows alongside cloud providers.
Generative-AI-enabled timeline video editor for macOS that lets creators generate and edit videos and images directly inside the timeline. Includes a local MCP server for agent integrations (Claude/Codex/Cursor); editor is open-source while generative processing is closed-source and subscription-based; macOS 26 on Apple Silicon only.
Automates video editing driven by LLM agents: reads word-level transcripts to propose and execute cuts, remove filler words, auto grade color, burn subtitles, and generate animation overlays. Self-evaluates every cut before showing a preview; aimed at talking-heads, tutorials and interviews.
Performs task-aware generative video restoration and editing in latent video space — restoration, super-resolution, watermark and subtitle removal — adapting LTX‑2.3 with IC‑Edit/IC‑LoRA adapters to prioritize temporal consistency and occlusion-aware reconstruction.
Enables Claude to “watch” videos by extracting timestamped frames plus captions/transcripts and feeding them to Claude for grounded Q&A. Key features: native captions first, Whisper fallback, frame deduplication, and multiple detail modes (transcript/efficient/balanced/token-burner). Useful for summarizing, debugging, and extracting moments.
Cross-platform native video editor with hardware-accelerated processing and frame-accurate multi-track timeline; core editor is open-source and free while optional Pro AI features (natural-language editing, auto-captions, smart reframing) are paid.
Converts video inputs into text outputs — supports captioning, temporal grounding, and video-text-to-text queries using a Qwen-3.5-2B finetuned multimodal backbone. Suited for prototyping video understanding and caption-generation pipelines.
Generates minute-scale, 720p videos from a single image using a 2.6B image-to-video diffusion transformer with precise 6‑DoF camera control and an optional LTX‑2 refiner; designed for long-context, memory-efficient modeling but requires large refiner checkpoints (~41 GB).
Generates temporally coherent MP4 videos from a single input image plus text instructions, with configurable resolution, frame count, and optional AAC audio. Optimized for NVIDIA GPU stacks and integrates with vLLM‑Omni and Hugging Face Diffusers for production inference and research workflows.
Generates audio-driven avatar videos from text, images, or audio inputs with production-grade stability (accurate lip sync, identity consistency) and an 8-step distillation inference mode for faster serving; suitable for broadcasting, virtual hosts, animation, and multi-person scenarios.