Reference architectures and microservices for building GPU-accelerated vision agents that enable natural-language video search, long-video summarization, visual Q&A, and alert verification. Integrates NVIDIA NIM models, embeddings, VLMs/LLMs, and agent workflows for deployable video-analytics stacks.
VideoCaptioner is an AI-powered video subtitling assistant that combines ASR (local or cloud) with LLM-based subtitle segmentation, correction and translation. It supports offline GPU transcription, concurrent chunk transcription, VAD, speaker-aware processing, batch subtitling and one-click subtitle-to-video synthesis, with both GUI and CLI options.
Generates video from text or images via a DiT-based latent diffusion model: text-to-video, image-to-video, frame extension, and multi-keyframe conditioning in one model. A distilled 2B variant runs near real-time on one H100; 13B for higher quality.
Custom ComfyUI nodes that run Lightricks' LTX-Video diffusion-transformer models for text-to-video and image-to-video, adding IC-LoRA control over depth, pose, edges, and motion plus distilled and low-VRAM variants for node-based workflows.
Retrieval-augmented generation framework for videos spanning hundreds of hours, runnable on a single RTX 3090. Builds multi-modal knowledge graphs over visual and audio content so you can query and chat across many long videos at once.
Distributes one post across 14+ platforms (Douyin, Xiaohongshu, TikTok, X), automates likes and replies via a browser plugin, and matches creators to paid brand tasks settled by sales, views, or engagement. Drivable from Claude/Cursor via MCP.
Provides curated ComfyUI workflow templates and subgraph blueprints that package reusable node graphs, preview assets, and publishing pipelines for image/video generation. Includes a browsable Astro site with i18n, CI-driven sync/publish scripts, and PyPI packaging for easy distribution.
Collects 40+ importable n8n workflows from the AI Agents A-Z YouTube channel, each tied to one video episode — spanning content generation, social-media posting, and short-video and narrated-story pipelines, plus companion Docker MCP/REST servers.
Turns a raw idea, novel, or screenplay into a complete multi-shot video through a multi-agent pipeline that scripts, storyboards, and renders shots while a vision model checks character and scene consistency across the whole story.
Clean-room, modular implementations of multi-object tracking algorithms — SORT, ByteTrack, OC-SORT, BoT-SORT, C-BIoU — behind one interface. Detector-agnostic: works with YOLO, DETR, or any bounding-box model via supervision.Detections.