AIAny
Icon for item

FineBooks BHL IMPACT Ground Truth

Provides 2,165 historical natural-history page scans paired with ~99.95% expert transcriptions and pixel-aligned PAGE XML layout ground truth for OCR and layout evaluation; multilingual (EN/FR/DE/LA), CC-BY 3.0.

Introduction

Why this matters

High-quality, page-level ground truth for historical print is scarce yet essential for developing and benchmarking OCR and document-layout models. This dataset supplies both near-perfect text transcriptions and precise polygonal layout annotations in the same pixel space as the shipped page images, enabling joint evaluation of text recognition and region detection on real historical material.

What Sets It Apart
  • Dual-purpose ground truth: each page pairs a ~99.95%-accurate reading-order transcription with full PAGE XML polygons, so you can evaluate OCR accuracy and layout/region detection without coordinate transforms.
  • Real historical diversity: 2,165 access pages drawn from six natural-history books (1708–1913) covering English, French, German and Latin (occasional Cyrillic in references), including running text, tables and illustrated plates.
  • Reproducible provenance: the ground truth originates from the IMPACT⇄BHL (2011–2012) collaboration and the dataset bundles the original GT XML, source scandata and the accessioned WebP access images used for mapping.
  • Production-ready packaging: images (WebP) + metadata.parquet + verbatim PAGE XML + derived markdown/docling exports make programmatic workflows (draw boxes, rebuild documents, join catalog metadata) straightforward.
Who it's for and trade-offs

Great fit if you need a moderate-sized, high-fidelity benchmark to train or evaluate OCR/text-recognition and document-layout models on historical print — especially for research on reading-order reconstruction, region detection, or multilingual historical transcription. The CC-BY 3.0 license simplifies reuse with attribution (IMPACT / BHL).

Look elsewhere if you need very large-scale corpora, non-access JP2 masters, or datasets with Gothic/Fraktur type — this corpus contains only antiqua types and includes access-page WebP renditions (≈390 MB). Note also that ~194 pages contain explicit U+FFFD tokens marking illegible glyphs; these are intentional unknown-character markers in the GT, not encoding errors.

Information

  • Websitehuggingface.co
  • Organizationsfinebooks (Hugging Face), IMPACT Centre of Competence, Biodiversity Heritage Library (BHL)
  • Published date2026/06/30

Categories

More Items

Hugging Face

Provides 5.5K+ self-contained data-analysis RL tasks: each row bundles a real tabular dataset, a question, and a deterministically-gradable gold answer. Verified from jupyter-agent notebooks; splits for training, held-out testing, and quick eval; intended for prompting, fine-tuning, and agent RL.

Hugging Face

Provides mixed-domain, verifiable RL training environments for LLM agents (code, cyber, knowledge work, web dev, music) as Parquet datasets, with domain-specific verifiers, Docker artifacts and links to training code for reproducible agentic RL experiments.

Provides WROP: a 1.5M-sample synthetic video corpus and a 300-question exam for training and evaluating object permanence and solidity in video world models. Includes 150 Blender task generators, a human Elo benchmark across 14 models, and a fine-tuned 16B continuation model (PWM-WROP).