Why this matters
High-quality, page-level ground truth for historical print is scarce yet essential for developing and benchmarking OCR and document-layout models. This dataset supplies both near-perfect text transcriptions and precise polygonal layout annotations in the same pixel space as the shipped page images, enabling joint evaluation of text recognition and region detection on real historical material.
What Sets It Apart
- Dual-purpose ground truth: each page pairs a ~99.95%-accurate reading-order transcription with full PAGE XML polygons, so you can evaluate OCR accuracy and layout/region detection without coordinate transforms.
- Real historical diversity: 2,165 access pages drawn from six natural-history books (1708–1913) covering English, French, German and Latin (occasional Cyrillic in references), including running text, tables and illustrated plates.
- Reproducible provenance: the ground truth originates from the IMPACT⇄BHL (2011–2012) collaboration and the dataset bundles the original GT XML, source scandata and the accessioned WebP access images used for mapping.
- Production-ready packaging: images (WebP) + metadata.parquet + verbatim PAGE XML + derived markdown/docling exports make programmatic workflows (draw boxes, rebuild documents, join catalog metadata) straightforward.
Who it's for and trade-offs
Great fit if you need a moderate-sized, high-fidelity benchmark to train or evaluate OCR/text-recognition and document-layout models on historical print — especially for research on reading-order reconstruction, region detection, or multilingual historical transcription. The CC-BY 3.0 license simplifies reuse with attribution (IMPACT / BHL).
Look elsewhere if you need very large-scale corpora, non-access JP2 masters, or datasets with Gothic/Fraktur type — this corpus contains only antiqua types and includes access-page WebP renditions (≈390 MB). Note also that ~194 pages contain explicit U+FFFD tokens marking illegible glyphs; these are intentional unknown-character markers in the GT, not encoding errors.