AIAny
Icon for item

UnsolvedMath

Provides a machine-readable collection of 5,426 open and historically significant mathematical problems with LaTeX statements, structured metadata and curated per-problem AI-assisted research notes. Includes difficulty labels, canonical problem sets (Millennium, Hilbert, Erdős) and files optimized for benchmarking math reasoning.

Introduction

Open mathematical problems function as a hard, non-memorisable stress test for AI: they require literature triage, formulation, partial progress and rigorous argumentation rather than single-answer recall. UnsolvedMath packages thousands of such problems with structured metadata and AI-assisted research audits so models can be evaluated on realistic research-style workflows rather than narrow QA tasks.

What Sets It Apart
  • Size and coverage: 5,426 problems spanning ~17 mathematical domains and multiple curated collections (Millennium Prize, Hilbert, Erdős, AMR lists). So what: gives benchmarks room to test across domain, difficulty, and problem-set distributions instead of a single-task slice.
  • Machine-readable LaTeX + structured fields: each problem includes a LaTeX statement, category, difficulty, provenance and status fields. So what: simplifies automated parsing, rendering and programmatic filtering for evaluation pipelines and tool-augmented research episodes.
  • AMR research audit and per-problem reports: thousands of AMR-triaged reports, per-problem research_classification and research_results files. So what: enables experiments that compare model-led literature triage, partial-progress generation, and citation-aware verification against a precomputed research baseline.
  • Designed for research workflows: dataset provides multiple JSON files (problems.json, research_results.json, categories, statistics) plus example loading snippets. So what: integrates easily with model evaluation, proof-checking workflows, and tool chains that combine symbolic/math systems and retrieval.
Who it's for + Tradeoffs

Great fit if you want to evaluate or develop models on genuine mathematical research tasks (literature search, proof sketching, counterexample finding, automated triage) across varied difficulty levels and curated problem sets. It is also useful for education and LaTeX-processing research.

Look elsewhere if you need a small, fully verified corpus of peer-reviewed theorem proofs for formal verification — the dataset prioritises breadth and machine-assisted triage over peer-reviewed closure. Note: status fields and AMR research notes can change; some SOLVED-* or PARTIAL-PROGRESS entries are machine-assisted and should be independently verified before citation.

Where It Fits

Compared with closed private problem collections used for automated numeric verification, UnsolvedMath emphasizes open problems, LaTeX fidelity and human-readable research reports aimed at evaluating end-to-end research behaviours (idea generation, literature triage, partial proofs) rather than single-answer scoring.

Information

Categories

More Items

Provides WROP: a 1.5M-sample synthetic video corpus and a 300-question exam for training and evaluating object permanence and solidity in video world models. Includes 150 Blender task generators, a human Elo benchmark across 14 models, and a fine-tuned 16B continuation model (PWM-WROP).

Hugging Face

Provides 1,800+ hours of synchronized egocentric multi-view recordings with 3D hand reconstructions, wide‑FOV depth, and hierarchical task/subtask annotations for embodied AI and robot learning. Includes six fisheye views, hand meshes, and per-episode temporal labels across 44k+ episodes.

Hugging Face

Provides 10,000 agentic multi-turn coding and reasoning traces from Fable 5.1 with step-by-step chain-of-thought and tool-use, heavily deduplicated and filtered; totals ~500M tokens (1.97 GB). Suited for SFT, distillation, and training long-horizon reasoning and code-generation models.