Large-scale synthetic video dataset of physically simulated multi-object interaction scenes for training and evaluating models on physical reasoning, depth and optical-flow estimation, instance segmentation, and physics-grounded captioning. Provides RGB + lossless depth, per-frame instance masks, per-object physics annotations (NPZ), VLM-grounded captions, and USD scene files — useful for world-model and simulation-to-real work; commercial use permitted.
Provides 10k–100k Indonesian-language cooking recipes in Parquet format, including dish names, ingredients and instructions — suitable for text-generation, recipe parsing, and culinary data analysis. Check the dataset card for license and field details.
Provides tick-aligned Counter-Strike 2 player POV video clips with per-tick inputs and world-state sidecars — near-lossless 1280×720@32fps video, per-player stereo audio, and parquet indexes for event/kill/round filtering; suited for RL, video classification and clip mining.
25,000 chat-formatted synthetic SFT examples distilled to emulate the reasoning style and agentic behavior of Anthropic's Claude Mythos, focused on cybersecurity, advanced coding, mathematical reasoning, and long-horizon agent tasks. Includes metadata for targeted curriculum fine-tuning and is Apache-2.0 licensed.
Provides 10,000 articulated 3D objects in URDF for robotics and embodied-AI research. Generated by the Articraft agent and released under CC-BY-4.0, the dataset targets simulation, manipulation, kinematics evaluation, and training of embodied agents.
Provides a 1-billion-parameter English pretrained language-model checkpoint that uses a dual-timescale Hierarchical Reasoning Model to increase effective compute depth. It's a PrefixLM pre-alignment checkpoint with composite-prefix modes for chain-of-thought style outputs; not instruction-tuned and requires downstream SFT/RL for assistant use.
Metadata-only corpus of 146.3M new GitHub source-code files (commit_id, rel_path, language) intended as an incremental update to Nemotron v1/v2 for LLM code pretraining; CC-BY-4.0 licensed and designed to be used jointly with older versions.
Automates distillation of heterogeneous traces from a target person or role into versioned, inspectable skill packages for LLM agents — producing separate capability and bounded-behavior tracks that support natural-language corrections, rollback, and cross-host installation. Ships with an open system and a skills gallery.
Provides ~1M synthetic Salvadoran‑Spanish personas (148k records, ~300M tokens) grounded in 2024 census distributions for demographics, occupations and locations; intended for training/evaluating localized LLMs and synthetic-data workflows. CC BY 4.0, adults only.
A surgically modified Gemma 4 (12B) that removes refusal behavior while preserving benchmark parity; released as an uncensored research artifact with GGUF quantizations for local inference and red‑team/alignment evaluation.
Pairs OCR-extracted annual-report text with ground-truth financial KPI values to benchmark LLM/table-QA and needle-in-a-haystack extraction tasks. Includes Markdown OCR (.mmd), page images for eval, and 31 KPI columns across multiple years—suited for KPI extraction, retrieval, and robustness testing.
Experimental, uncensored fine-tune of Google Gemma-4-12B-it that applies an 'abliteration' technique to remove refusal behaviors; intended for research and testing only and carries elevated safety and legal risks.