Multimodal STEM problem set for verifiable, answer-supervised training and RL: contains single-image, multi-panel, and multi-image PhD-level questions across physics, math, chemistry and biology. Each example has a deterministic ground-truth answer, enabling reward modeling and automated evaluation.
Pairs natural-language instructions with executable setup artifacts and Python reward functions to create verifiable computer-use agent tasks. Provides a Parquet task table for fast filtering plus a compressed archive of runnable task bundles; several web task endpoints are placeholders that require a local CUA-Gym-Hub deployment.
Provides 40 public Kubernetes incident scenarios (SRE subset) with ground-truth root-cause entities and offline cluster snapshots in JSONL format; designed to evaluate agentic root-cause diagnosis on alerts, events, traces and topology.
Provides 9,000 reconstructed chain-of-thought (CoT) SFT examples produced by trace inversion from Claude Opus 4.6 outputs for fine-tuning reasoning-capable LLMs. Multilingual, packaged as .jsonl.gz and SFT/DPO-ready; verify numeric/code cases before training.
Parallel Chinese→Vietnamese dataset of webnovel (xianxia) text provided in JSON for NMT training and teacher-student distillation. In-domain, ~100K–1M examples with CC-BY-4.0 license — useful for fine-tuning or distillation experiments but limited by narrow genre and small download footprint.
A human‑curated corpus of AI‑generated music with MP3s, cover art, exact generation prompts and a 32‑column metadata schema; uses a 70/30 quality vs. mainstream split and a three‑level taxonomy to support fine‑grained audio‑ML, prompt‑fidelity and recommendation research.
Metadata-only corpus of 146.3M new GitHub source-code files (commit_id, rel_path, language) intended as an incremental update to Nemotron v1/v2 for LLM code pretraining; CC-BY-4.0 licensed and designed to be used jointly with older versions.
A collection of 14,056 self-contained research-level mathematical problems extracted from papers and open-problem lists, each rewritten with taxonomy labels and open-status metadata for training or evaluating models on research-grade math reasoning.
Provides a sanitized, MIT‑licensed dataset of scanner evidence and registry verdicts for public ClawHub agent skills — 67k+ latest skill versions with redacted artifacts and structured VirusTotal, static-analysis, and SkillSpector outputs to study scanner disagreement and agent-skill risk governance.
Evaluates metric 3D spatial reasoning from single driving images via multiple-choice questions that require reconstructing scene geometry rather than relying on image-layout shortcuts. Each sample pairs a numbered-bbox image with a question, four choices, and the correct answer; images come from PlusAI and the dataset is CC BY 4.0.
A 16 GB, 507-file PhD‑level cybersecurity knowledge base for training and evaluating security-focused LLMs and automation. Covers offensive/defensive/forensics/cloud/iot and AI-security across 30+ domains with real-world labs and framework mappings.
Provides the gated, official OSWorld 2.0 Python task class files (task_*.py) required to run the benchmark; distributed via a Hugging Face gated dataset to reduce benchmark leakage. Download requires accepting gated access on Hugging Face.