Historical OCR archives often separate text dumps from the page images or omit pixel-accurate layout. This dataset joins full IIIF-served page images with the original ALTO XML converted into image-pixel line and word boxes plus per-word confidence scores, making it practical to train and evaluate line-level OCR, spot bad regions for re-OCR, and train layout-reading models across scripts and languages.
What Sets It Apart
- Pixel-accurate OCR layout: every ALTO TextLine and String is converted to image-pixel bounding boxes (lines[].bbox, words[].bbox) so you can crop exact line images or compute per-word losses. This removes ad-hoc coordinate transforms for most pages.
- Rich, silver-quality targets at scale: 98,877 pages across 10 languages and multiple scripts (Fraktur, Cyrillic, Greek), with ALTO v2 XML, mean_ocr and per-word wc values — large enough for pretraining, domain adaptation, and systematic error analysis, but not a gold-standard transcription corpus.
- IIIF-first images: images are the original IIIF-served scans (varying resolutions by holding library); the dataset includes page_iiif_url and image width/height so you can stream or fetch only what you need.
- Sampling + metadata workflow: a small metadata config lets you pick pages before downloading large rows; many workflows stream images and avoid loading everything into RAM.
Who it's for
Great fit if you want realistic, historical OCR training or evaluation data (line-level supervision, confidence-aware filtering, layout and reading-order models) or if you need paired image+ALTO to compare modern OCR/VLM outputs against legacy OCR. Look elsewhere if you need gold-standard transcriptions for evaluation, a fully representative sample of Europeana, or consistently high-resolution images — OCR quality and image sizes vary by collection, and some pages have alignment issues marked in box_alignment.
Where it fits
Use this as silver-scale data for pretraining, error-finding and re-OCR pipelines, or as a diverse visual+text sample for document-understanding model development; complement with smaller hand-corrected datasets for final benchmarks.