AIAny
Icon for item

openai/math as Harbor environments

Provides 369 Harbor sandbox tasks ported from OpenAI's openai/math: each task is a Lean theorem with missing `sorry` proofs that an agent must complete, graded by a strict Comparator exact-match verifier. Includes task definitions, generator, and manifest for RL/code-agent evaluation.

Introduction

These Harbor environments turn OpenAI's formalized Lean theorems into reproducible, RL-style proof-generation tasks where an agent must replace every sorry with a complete Lean proof and pass an exact-match verifier.

What Sets It Apart
  • Directly reuses OpenAI's openai/math formalizations but packages them as 369 verified Harbor tasks (from 405 formalized results). Each task bundles the Lean statement, a sandbox image recipe (Lean 4 + Mathlib + Comparator), and a strict verifier that rewards only exact, kernel-accepted proofs. The dataset includes a manifest, task metadata, and a generator for rebuilding images.
  • Designed for code agents and theorem-proving research: tasks allocate CPU/memory/disk (typical task: 4 CPUs, 8 GB RAM, 10 GB disk, 4 hours) and the oracle can reproduce OpenAI's proofs for validation. The grading is binary and narrow by design—only the exact original statement proved with no new axioms scores positive reward.
Who It's For and Trade-offs
  • Great fit if you build or benchmark code agents, autonomous provers, or RL systems that generate formal Lean proofs and need a reproducible, high-fidelity verifier. Also useful for studying reward design in long-horizon, symbolic tasks.
  • Look elsewhere if you need permissive grading, partial-credit evaluation, or larger-scale formalizations: the verifier is intentionally strict (exact-statement match) and some proofs require >8 GB memory or external Lean libraries and so are excluded or flagged. Expect hard, research-level problems rather than toy tasks.

Information

  • Websitehuggingface.co
  • OrganizationsFineEnvs, OpenAI, Harbor, leanprover-community, Comparator
  • Published date2026/10/08

Categories

More Items

Hugging Face

Provides a 1 trillion-token multimodal interleaved dataset (HTML subset updated as data_v1_1 with 742B HTML tokens) and 3.4B images drawn from HTML/PDF/ArXiv sources for multimodal pretraining; released under CC-BY-4.0 with safety and deduplication guidance.

Hugging Face

Curated English Wikipedia text prepared for language-model training and evaluation, provided in WikiText-2 and WikiText-103 variants. Preserves original case, punctuation and numbers; offers raw and tokenized splits for long-range language modeling under a CC BY‑SA license.

Hugging Face

Provides 16 weeks of anonymized production agent-session traces (12,002 sessions, ~1.19M LLM requests, ~1.21M tool calls, 209B input tokens) released as block-level prefix IDs plus flattened Parquet tables for KV-cache, scheduling and serving-system research.