Delivers an ultra-efficient, edge-friendly multimodal image-and-video-to-text model optimized for on-device deployment. Uses mixed 4x/16x visual token compression, a low-FLOPs visual encoder, and multiple quantized variants for mobile and embedded inference.
Reconstructs camera poses and dense 3D point clouds from video streams using a feed‑forward foundation model. Combines a Geometric Context Transformer (anchor + local window + trajectory memory) with paged KV‑cache attention to enable stable, long‑sequence streaming inference (~20 FPS at 518×378).
Turns books, long videos, and podcasts into executable, testable AI agent skills using a structured RIA‑TV++ pipeline. Produces multi-file skill packs (BOOK_OVERVIEW.md, SKILL.md, INDEX.md, DIGEST.md), applies triple verification and pressure tests, and can install skills into Claude Code/Cursor for agent use.
Unified multimodal LLM for enterprise workflows: ingests video, audio, image and text to perform transcription, OCR, Q&A, summarization and long-context reasoning. Provides BF16/FP8/NVFP4 weights and integrations with vLLM, TensorRT-LLM and other runtimes.
An HDR LoRA fine-tune for Lightricks' LTX-2.3 (22B) that enables image‑conditioned any‑to‑any image-to-video and text-to-video generation. Designed for HDR-aware synthesis workflows; requires the LTX-2.3 base model and a LoRA-capable runtime.
Performs task-aware generative video restoration and editing in latent video space — restoration, super-resolution, watermark and subtitle removal — adapting LTX‑2.3 with IC‑Edit/IC‑LoRA adapters to prioritize temporal consistency and occlusion-aware reconstruction.
Unifies video, audio, image and text understanding for enterprise Q&A, summarization, transcription and document intelligence. The NVFP4 quantized variant reduces footprint to ~20.9GB for more efficient single‑GPU deployment and is tuned for NVIDIA runtimes (vLLM, TensorRT).
Cross-platform native video editor with hardware-accelerated processing and frame-accurate multi-track timeline; core editor is open-source and free while optional Pro AI features (natural-language editing, auto-captions, smart reframing) are paid.
Unified omnimodal foundation model for text, image, video and audio understanding and agentic workflows, with support for up to 1M-token context. Combines a sparse MoE LLM backbone, dedicated vision/audio encoders, multi-token prediction, and a hybrid sliding-window + global attention design to reduce KV-cache overhead.
Provides a GGUF-quantized build of NVIDIA's Nemotron 3 Nano Omni 30B (Reasoning) for local inference — enables multimodal (video/audio/image/text) reasoning, transcription, and document understanding on compatible runtimes such as llama.cpp, Ollama, vLLM, and TensorRT-LLM.
Large-scale synthetic video dataset of 236,937 1080p clips (≈5,841 hours) of digital humans with per-frame metric depth and camera parameters — built as a controllable supplement for world-model pretraining, camera-motion generalization, and geometry-aware physical-AI research.
Structured, downloadable JSONL dataset of Seedance 2.0 video-generation prompts with matching MP4 previews and cover images; includes English/Chinese texts and standardized metadata (duration, resolution, safety) and is released under CC BY 4.0 for reuse.