AIAny
Icon for item

FineBooks BHL IMPACT Ground Truth

Provides 2,165 historical natural-history page scans paired with ~99.95% expert transcriptions and pixel-aligned PAGE XML layout ground truth for OCR and layout evaluation; multilingual (EN/FR/DE/LA), CC-BY 3.0.

Introduction

Why this matters

High-quality, page-level ground truth for historical print is scarce yet essential for developing and benchmarking OCR and document-layout models. This dataset supplies both near-perfect text transcriptions and precise polygonal layout annotations in the same pixel space as the shipped page images, enabling joint evaluation of text recognition and region detection on real historical material.

What Sets It Apart
  • Dual-purpose ground truth: each page pairs a ~99.95%-accurate reading-order transcription with full PAGE XML polygons, so you can evaluate OCR accuracy and layout/region detection without coordinate transforms.
  • Real historical diversity: 2,165 access pages drawn from six natural-history books (1708–1913) covering English, French, German and Latin (occasional Cyrillic in references), including running text, tables and illustrated plates.
  • Reproducible provenance: the ground truth originates from the IMPACT⇄BHL (2011–2012) collaboration and the dataset bundles the original GT XML, source scandata and the accessioned WebP access images used for mapping.
  • Production-ready packaging: images (WebP) + metadata.parquet + verbatim PAGE XML + derived markdown/docling exports make programmatic workflows (draw boxes, rebuild documents, join catalog metadata) straightforward.
Who it's for and trade-offs

Great fit if you need a moderate-sized, high-fidelity benchmark to train or evaluate OCR/text-recognition and document-layout models on historical print — especially for research on reading-order reconstruction, region detection, or multilingual historical transcription. The CC-BY 3.0 license simplifies reuse with attribution (IMPACT / BHL).

Look elsewhere if you need very large-scale corpora, non-access JP2 masters, or datasets with Gothic/Fraktur type — this corpus contains only antiqua types and includes access-page WebP renditions (≈390 MB). Note also that ~194 pages contain explicit U+FFFD tokens marking illegible glyphs; these are intentional unknown-character markers in the GT, not encoding errors.

Information

  • Websitehuggingface.co
  • Organizationsfinebooks (Hugging Face), IMPACT Centre of Competence, Biodiversity Heritage Library (BHL)
  • Published date2026/06/30

Categories

More Items

Hugging Face

Structured, downloadable JSONL dataset of Seedance 2.0 video-generation prompts with matching MP4 previews and cover images; includes English/Chinese texts and standardized metadata (duration, resolution, safety) and is released under CC BY 4.0 for reuse.

Hugging Face

Provides per-decision training samples for RL-driven command-line LLM agents: each record pairs a task prompt plus terminal history with a teacher's next-action in Terminus-2 JSON. Around 31k verifier-passing samples from 630 ATCB tasks, formatted for NeMo Gym's terminus_judge and licensed CC-BY-4.0.

Hugging Face

Benchmark for joint speaker diarization and speaker-attributed ASR across all 22 scheduled Indian languages, providing ~108 hours of human-corrected, time-aligned, speaker-attributed transcripts. Includes near-field, far-field and in-the-wild recordings with code-mixing and speaker overlap.