Tag
Explore by tags
Author HTML-based video compositions and render deterministic, frame-accurate MP4s with agent-friendly tooling — preview in the browser, drive generation via AI agent skills, and use adapter runtimes (GSAP, Lottie, Three.js).
Provides an annotated multimodal human-motion dataset for language-to-action and robotics research, with BVH and MuJoCo files plus recordings targeted at Unitree-G1 and NVIDIA-SOMA platforms. Covers locomotion, gestures, dance and object interaction with English annotations and 100K–1M samples.
Generate text, images, video, audio and action/robot trajectories from combined text, image, video, audio and action inputs. A Mixture-of-Transformers omnimodal foundation model (Cosmos3‑Nano, 16B params) focused on Physical AI (robotics, AV, simulation) and optimized for NVIDIA GPU runtimes.
Provides 10 million synchronized egocentric experience episodes with structured 3D/4D multimodal annotations — 2.88B RGB frames, 720M depth frames, 576M pose/mocap frames and ~1PB total. Designed for embodied AI, robotics, and multimodal pretraining; research-only, gated access.
Large-scale mid-training corpora for multimodal models: 10,809 ~60s video shards, caption splits (30s/60s/180s/>10min), 84 spatial-reasoning shards, and CSV mappings to source YouTube IDs. Small Parquet preview configs are provided for schema inspection.
Orchestrates end-to-end video production with agentic pipelines that research, script, generate assets, edit, and render finished videos. Distinguishes itself by supporting true real-footage retrieval (Archive.org, NASA, Wikimedia), Remotion/HyperFrames composition, and usable zero-key workflows alongside cloud providers.
Provides a diagnostic suite that audits video-understanding benchmarks to find samples solvable without visual or temporal input, filters those shortcuts, and produces a distilled video-native testbed that reveals major capability gaps in current Video-LLMs.
Provides curated short video clips (49- and 81-frame) with layered ground truth—edit layers, alpha mattes, and composite targets—for training and evaluating content-preserving layered diffusion video editing. Contains background-replace and object-add edits; Apache-2.0 licensed.
Generative-AI-enabled timeline video editor for macOS that lets creators generate and edit videos and images directly inside the timeline. Includes a local MCP server for agent integrations (Claude/Codex/Cursor); editor is open-source while generative processing is closed-source and subscription-based; macOS 26 on Apple Silicon only.
Desktop app for local voice cloning, real-time dictation, and end-to-end video dubbing using zero-shot TTS across 600+ languages; features multi-engine TTS/ASR, speaker diarization, vocal isolation, batch pipelines, and invisible audio watermarking — all run fully offline.
Large-scale in-the-wild robot manipulation dataset with ~76K teleoperated trajectories (~350 hours) that provides synchronized multi-view video, depth, camera calibration, robot state/action traces, and natural-language task instructions to train and evaluate manipulation policies and dynamics models. Collected across 564 scenes, 86 tasks, 52 buildings, on a uniform Franka Panda hardware stack and released in LeRobotDataset v3.0 format (≈707 GB, OpenMDW1.1).
Automates video editing driven by LLM agents: reads word-level transcripts to propose and execute cuts, remove filler words, auto grade color, burn subtitles, and generate animation overlays. Self-evaluates every cut before showing a preview; aimed at talking-heads, tutorials and interviews.