Provides labeled movie-review data for binary sentiment classification: 25,000 training and 25,000 test examples, plus 50,000 unlabeled reviews for unsupervised or semi-supervised use. Labels reflect strong polarity (positive ≥7, negative ≤4) and the set is a widely used NLP benchmark.
Provides about 100,000 crowd‑written question–answer pairs from Wikipedia where each answer is a text span in the passage, used to train and evaluate extractive question‑answering models. Includes train/validation splits, span offsets, Parquet format, CC BY‑SA 4.0.
A multi-task English NLU benchmark for evaluating models across nine tasks (acceptability, sentiment, paraphrase, similarity, and various NLI setups), with a diagnostic evaluation set and an online leaderboard to compare generalization and transfer learning.
Provides manually curated Japanese instruction pairs (questions and safe reference answers) for improving LLM output safety, covering broad harm categories and regionally sensitive cases. Includes English meta-tags and standard splits for benchmarking and fine-tuning.
Provides low‑latency on‑device speech-to-text, intent recognition, and text-to-speech for building real‑time voice agents and interfaces. Streaming-optimized models, incremental caching, multilingual TTS/ASR and cross-platform bindings (Python, iOS, Android, Linux, Raspberry Pi) target live voice use cases where sub-200ms responsiveness matters.
Measures generative AI inference performance with token-level metrics (TTFT, inter-token latency), latency, and throughput under realistic traffic patterns. Provides a multiprocess engine, real-time TUI dashboard, extensible plugins, and integrations for telemetry and result uploads, aimed at inference benchmarking and capacity planning.
Provides 1.7M+ synthetic and real infographic charts paired with their tabular data for training and evaluating multimodal models on infographic understanding, chart-to-table extraction, chart code generation, and example-based chart synthesis.
Provides MS MARCO queries, passages and answers translated into 14 Indic languages while keeping the original English content and per-example translation metadata. Includes train/validation splits, passage selection flags, and translation model parameters for multilingual IR, QA and RAG research.
Evaluates and optimizes AI agents and language models in containerized environments, supporting large-scale parallel benchmarks and RL rollouts. Integrates with third‑party providers for thousands of parallel environments and serves as the official harness for Terminal‑Bench.
Open-source companion to a technical book that teaches how to design, evaluate and ship LLM-based AI agents — includes the full Chinese manuscript, community translations, chapter-aligned runnable example projects, and reproducible evaluation harnesses.
Provides a machine-readable collection of 5,426 open and historically significant mathematical problems with LaTeX statements, structured metadata and curated per-problem AI-assisted research notes. Includes difficulty labels, canonical problem sets (Millennium, Hilbert, Erdős) and files optimized for benchmarking math reasoning.
Provides a physical reconstruction benchmark of OmniDocBench v1.5 by producing five real-world photographic variants (Scanning, Warping, Screen‑Photography, Illumination, Skew) for each of 1,355 pages, inheriting original ground-truth to enable controlled, scenario-wise evaluation of document parsing robustness.