Provides 1.7M+ synthetic and real infographic charts paired with their tabular data for training and evaluating multimodal models on infographic understanding, chart-to-table extraction, chart code generation, and example-based chart synthesis.
Provides MS MARCO queries, passages and answers translated into 14 Indic languages while keeping the original English content and per-example translation metadata. Includes train/validation splits, passage selection flags, and translation model parameters for multilingual IR, QA and RAG research.
Provides a large-scale, multi-speaker Persian speech–text corpus constructed from audiobooks for TTS, ASR, and speaker research. Includes automated alignment and quality scoring, TTS-ready subsets (thousands of hours/1M+ segments) and metadata for speaker IDs and genders — suitable for multi-speaker synthesis and voice cloning research.
Provides a reproducible, deduplicated corpus of text extracted from PDFs for LLM pretraining—about 3 trillion tokens from ~475 million documents in 1733 language-script pairs. Includes OCR and text extraction pipelines, per-page language IDs, MinHash deduplication, and is released under ODC‑By 1.0.
Provides 6,000 runnable, operator-level PyTorch tasks for training and evaluating CUDA kernel generation models; each sample includes executable code, operator descriptors, and provenance tags, with execution-driven filtering to ensure reproducibility and contamination control.
Provides 3,000+ hours (≈611K utterances) of transcribed 16 kHz multi-dialect Arabic speech across 13 dialects for ASR and spoken-dialect identification. Transcripts preserve dialectal orthography (partial diacritics); the train split is ~337 GB in parquet, so streaming is recommended.
Large-scale in-the-wild robot manipulation dataset with ~76K teleoperated trajectories (~350 hours) that provides synchronized multi-view video, depth, camera calibration, robot state/action traces, and natural-language task instructions to train and evaluate manipulation policies and dynamics models. Collected across 564 scenes, 86 tasks, 52 buildings, on a uniform Franka Panda hardware stack and released in LeRobotDataset v3.0 format (≈707 GB, OpenMDW1.1).
Provides 545,431 math problems with model-generated solution traces (chain-of-thought and Python tool-integrated reasoning) verified against reference answers for training and evaluating LLM mathematical reasoning. Parquet-format dataset; DeepSeek‑V4‑Pro generated traces and mixed CC BY / CC BY‑SA licensing.
Provides 1.3 billion platform-specific video URLs extracted from CommonCrawl along with crawl metadata (no media included), serving as the source corpus for the LAION-BVD multimodal video dataset; distributed on Hugging Face in Parquet format.
A 16 GB, 507-file PhD‑level cybersecurity knowledge base for training and evaluating security-focused LLMs and automation. Covers offensive/defensive/forensics/cloud/iot and AI-security across 30+ domains with real-world labs and framework mappings.
Benchmark dataset for evaluating long-horizon coding agents and software-engineering tasks, containing English code and tabular metadata in Parquet format; small scale (<1K examples) for fast prototyping and evaluation.
Self‑hosted A‑share quantitative workbench for screening, monitoring, backtesting and stock-level analysis using TickFlow data; supports 18 Polars strategies, vectorbt backtesting, realtime rule-based alerts, pluginable data sources and optional LLM-driven strategy generation and stock analysis.