Provides a county-harmonized corpus of U.S. municipal and county ordinance text (≈2.21M chunks) labeled for function, substantive indicator, and topic to support legal NLP, retrieval, and comparative local-law research. Includes model-assigned labels and continuous scorers (opacity, paternalism, enforcement discretion) plus coverage metadata; not exhaustive or a substitute for legal advice.
Provides 40 public Kubernetes incident scenarios (SRE subset) with ground-truth root-cause entities and offline cluster snapshots in JSONL format; designed to evaluate agentic root-cause diagnosis on alerts, events, traces and topology.
Provides per-cell transcriptomes and five-day drug-sensitivity readouts for 1.83M single cells across 52 cancer cell lines and 91 drug conditions, with raw counts plus gene, cell-line, drug, and summary metadata for modeling drug response and context-dependent gene function.
Provides the gated, official OSWorld 2.0 Python task class files (task_*.py) required to run the benchmark; distributed via a Hugging Face gated dataset to reduce benchmark leakage. Download requires accepting gated access on Hugging Face.
Benchmark dataset for evaluating long-horizon coding agents and software-engineering tasks, containing English code and tabular metadata in Parquet format; small scale (<1K examples) for fast prototyping and evaluation.
Provides 500+ hours of human whole-body teleoperation recordings of a Unitree G1 in real homes, packaged in LeRobot v3.0 for robot learning. Contains 23K+ episodes, ~40M frames, multi-view 480p@30 video, 29-DoF states, actions and language annotations; CC BY 4.0 and large download size.
Contains a sanitized Claude Code (Fable 5) JSONL transcript of a session that procedurally built a Boeing 747 in Three.js, including assistant messages, tool calls, and base64 screenshots — useful for studying agent trace, tool use, and vision self‑verification workflows.
Refines large-scale English pretraining corpora by predicting per-instance structured edits (insert, delete, replace) and deterministically applying them to produce cleaner text for LLM training. Provides five ~20B-token refined corpora in parquet with edit metadata and simple loading configs.
Benchmark for evaluating multimodal LLM safety in Korean cultural contexts — includes KSAFE-MM-G which localizes global safety queries into Korean scenarios and KSAFE-MM-C which targets culture-specific visual-textual vulnerabilities. Provides curated image–text pairs and jailbreak-style prompts to reveal both unsafe behaviors and over-refusal.
Provides kanji-level evaluation data for Japanese TTS: disambiguated sentence contexts targeting 4,378 kanji-reading pairs (2,136 Jōyō kanji) with 13,095 native-speaker–verified sentences and katakana-marked ground-truth readings for kanji-level error metrics.
Collects raw coding-agent sessions—developer prompts, model replies, tool calls, and command output—donated from public repositories and anonymized locally. Organized by agent harness (raw session files + Parquet table), useful for studying agent behavior and tool use; anonymization is best-effort.
A JSON dataset of ~1.1M anonymized coding-assistant instruction→response interactions for training and evaluating code-generation and instruction-following models; packaged for use with pandas/polars and sized at ~459 MB.