Real‑time full‑duplex speech‑to‑speech system that controls conversational role via text prompts and voice timbre via audio-conditioned embeddings. Built on Moshi; optimized for low-latency, persona-consistent spoken interactions.
Large-scale mathematical reasoning dataset of model-generated solution trajectories produced with and without Python Tool-Integrated Reasoning (TIR), with final answers verified against reference solutions. Contains ~3.64M JSONL training samples (~144 GB) and per-source CC-BY / CC-BY-SA licensing; intended for training and evaluating tool-augmented mathematical reasoning in LLMs.
A transparent proxy that lets Claude Code CLI and the VSCode extension run without an Anthropic API key by routing requests to NVIDIA NIM, OpenRouter, LM Studio, or llama.cpp; supports per-model mapping, thinking-token parsing, and messaging integration.
Generates anime-style and other non-photorealistic illustrations from text prompts. A 2B-parameter diffusion base preview trained on millions of anime images (and ~800k non-anime art) and released under a non-commercial license; best used in ComfyUI around ~1MP resolution.
Provides a catalog of NVIDIA-verified, portable “skills” — instruction sets that teach AI agents how to use NVIDIA libraries, models and platform tools. Each skill is published with detached signatures and evaluation artifacts for verifiable reuse in agent workflows.
Provides an annotated multimodal human-motion dataset for language-to-action and robotics research, with BVH and MuJoCo files plus recordings targeted at Unitree-G1 and NVIDIA-SOMA platforms. Covers locomotion, gestures, dance and object interaction with English annotations and 100K–1M samples.
Generate text, images, video, audio and action/robot trajectories from combined text, image, video, audio and action inputs. A Mixture-of-Transformers omnimodal foundation model (Cosmos3‑Nano, 16B params) focused on Physical AI (robotics, AV, simulation) and optimized for NVIDIA GPU runtimes.
Scans AI agent skills for security issues—detecting vulnerabilities, malicious patterns, and supply-chain risks before installation. Combines static AST checks (64 patterns across 16 categories) with optional LLM semantic review, OSV live CVE lookups, and JSON/Markdown/SARIF outputs for CI or manual review.
Generates persistent, explorable 3D worlds from a single image by synthesizing long-range, geometry-consistent video and reconstructing it into an explicit 3D Gaussian scene. Intended for internal research use under NVIDIA's research license.
Provides 12.26M synthetically generated multilingual OCR samples (en/ja/ko/ru/zh) with word/line/paragraph bounding boxes and reading-order graphs, packaged as HDF5 shards for training detection, recognition, and layout models; licensed CC BY 4.0.
Generates text by iteratively denoising blocks of tokens with a two-tower design: a frozen autoregressive context tower and a trainable diffusion denoiser tower, trading minimal quality loss for higher wall-clock throughput.
Large-scale in-the-wild robot manipulation dataset with ~76K teleoperated trajectories (~350 hours) that provides synchronized multi-view video, depth, camera calibration, robot state/action traces, and natural-language task instructions to train and evaluate manipulation policies and dynamics models. Collected across 564 scenes, 86 tasks, 52 buildings, on a uniform Franka Panda hardware stack and released in LeRobotDataset v3.0 format (≈707 GB, OpenMDW1.1).