Provides 1,000,000 model-generated chain-of-thought traces and instruction–response pairs for fine-tuning and distilled supervision. Focused splits (coding, PHD-Science, General-Math, MultilingualSTEM), ~5B tokens, Apache-2.0 license.
Provides a unified 615k-hour English speech corpus for TTS training, aggregating 11 public datasets and web-sourced recordings into 16 kHz Opus WebDataset shards. Includes a quality-filtered core subset (510.1k hours), metadata splits, and mixed licenses across sources.
Provides 3,000+ hours (≈611K utterances) of transcribed 16 kHz multi-dialect Arabic speech across 13 dialects for ASR and spoken-dialect identification. Transcripts preserve dialectal orthography (partial diacritics); the train split is ~337 GB in parquet, so streaming is recommended.
Provides multi-turn agent trajectories with real tool executions and explicit <think> reasoning blocks for training and evaluating tool-calling agents. Contains two model-sourced configs (Kimi-K2.5, GLM-5.1) totaling ~14.7K samples — useful for SFT, agent-skill research, and tool-integration experiments.
Benchmarks LLM agents on realistic legal work by packaging lawyer-style assignments with client materials and expert, per-deliverable rubrics. Includes an execution harness to run, score, and compare agents across a large, evolving task set spanning multiple practice areas.
Aggregates and deduplicates public Claude distillation datasets into a unified 'messages' format with source attribution; focused on instruction-tuning and reasoning samples for SFT and LLM training, while requiring users to follow original sources' licenses.
Provides curated short video clips (49- and 81-frame) with layered ground truth—edit layers, alpha mattes, and composite targets—for training and evaluating content-preserving layered diffusion video editing. Contains background-replace and object-add edits; Apache-2.0 licensed.
Provides 12.26M synthetically generated multilingual OCR samples (en/ja/ko/ru/zh) with word/line/paragraph bounding boxes and reading-order graphs, packaged as HDF5 shards for training detection, recognition, and layout models; licensed CC BY 4.0.
Benchmarks document-parsing systems on real-world enterprise PDFs and images—evaluates tables, charts, content faithfulness, semantic formatting, and visual grounding with human-verified, rule-level tests. Ships with ~2,000 pages, ~169K test rules, and an open evaluation framework for end-to-end pipeline scoring.
Curated 100K subset of geometrically diverse CAD construction sequences sampled from a 1M agentically synthesized corpus — each item includes executable CadQuery scripts, 8 rendered views, STL/STEP exports, and precomputed DINOv3 embeddings for retrieval and benchmarking.
Provides a ~9.2M-instance Japanese multimodal post-training dataset for vision–language models, combining image–text pairs, PDF corpora and generated VQA to improve Japanese VLM performance; access is restricted by Japanese copyright (download via llm-jp GitLab).
Provides one million executable, human-readable CadQuery construction sequences synthesized by an LLM-in-the-loop—each sample includes renders, STL/STEP exports, precomputed DINOv3 embeddings and a FAISS index. Designed for training and benchmarking text/image→3D and CAD-program generation models (Apache-2.0).