Provides 5 million instruction–response pairs for supervised fine-tuning of code LLMs, with inputs, outputs, unit tests, and automated LLM judgments. Uses hybrid automated/synthetic generation and is released under CC BY 4.0 for large-scale SFT workflows.
Measures generative AI inference performance with token-level metrics (TTFT, inter-token latency), latency, and throughput under realistic traffic patterns. Provides a multiprocess engine, real-time TUI dashboard, extensible plugins, and integrations for telemetry and result uploads, aimed at inference benchmarking and capacity planning.
GPU-accelerated physics simulation engine for robotics and simulation research — built on NVIDIA Warp with MuJoCo Warp backend, offering differentiable simulation, OpenUSD support, and extensions for RL/embodied-AI workflows. ([github.com](https://github.com/newton-physics/newton))
A PyTorch DTensor-native SPMD library for training and fine-tuning LLMs, VLMs, diffusion and retrieval models. Integrates with Hugging Face for day-0 model support, provides YAML-driven recipes, DTensor/FSDP2 parallelism and NVIDIA-optimized kernels (Transformer Engine, DeepEP, FlexAttn).
Enables bidirectional checkpoint conversion between Hugging Face and Megatron formats and provides a PyTorch-native training library with tensor/pipeline parallelism, FP8/BF16 mixed precision, SFT and PEFT (LoRA) support for large and multimodal models.
Physics-aware simulated sensor dataset for training and evaluating autonomous-vehicle perception and control models. Includes multimodal sensor streams with physical-scene annotations intended for tasks that require grounding in real-world dynamics.
1,000,000 US-focused synthetic persona records (6M persona texts) grounded to demographic, geographic and personality distributions. Contains age, sex, education, occupation and ZCTA/city fields; CC BY 4.0 license for LLM training, diversity augmentation, and bias mitigation.
Generates explorable, 3D-consistent virtual worlds from a single image or short video. Includes official implementations of Lyra‑1 (feed‑forward 3D/4D scene generation via video-diffusion self-distillation) and Lyra‑2 (long-horizon, explorable generative 3D worlds). Best for research and creative prototyping; requires substantial GPU compute.
Provides an NVFP4‑optimized training and inference infrastructure for long-form video diffusion models — supports multi-shot AR training, KV-cache and NVFP4 quantized inference, sequence-parallelism and async decoding for higher FPS and longer outputs.
Generates production-grade synthetic datasets from scratch or from seed data using dependency-aware samplers, LLM-backed text columns, built-in validators, previewing, and LLM-as-judge scoring.
Contains training, evaluation, and deployment code plus checkpoints for humanoid whole-body controllers (Decoupled WBC and GEAR‑SONIC). Includes C++ inference, VR teleoperation, data pipelines (Bones‑SEED) and Hugging Face checkpoints for research-to-robot workflows.
Converts images (and other conditions) into high-fidelity, fully textured 3D assets using a 4B-parameter generative model and a field‑free sparse voxel format (O‑Voxel). Handles arbitrary topology, PBR materials, and near real-time mesh/voxel conversions; requires Linux and an NVIDIA GPU with >=24GB memory.