AIAny
Icon for item

Europeana Newspapers: Page Images with OCR Layout

Provides 98,877 historical newspaper page images (1700s–1940s) paired with ALTO OCR, per-word confidences and line/word bounding boxes in image pixels — ready for line-level OCR training, OCR quality estimation, re‑OCR comparisons and layout analysis. OCR is library-produced (silver); image resolutions and OCR quality vary.

Introduction

Historical OCR archives often separate text dumps from the page images or omit pixel-accurate layout. This dataset joins full IIIF-served page images with the original ALTO XML converted into image-pixel line and word boxes plus per-word confidence scores, making it practical to train and evaluate line-level OCR, spot bad regions for re-OCR, and train layout-reading models across scripts and languages.

What Sets It Apart
  • Pixel-accurate OCR layout: every ALTO TextLine and String is converted to image-pixel bounding boxes (lines[].bbox, words[].bbox) so you can crop exact line images or compute per-word losses. This removes ad-hoc coordinate transforms for most pages.
  • Rich, silver-quality targets at scale: 98,877 pages across 10 languages and multiple scripts (Fraktur, Cyrillic, Greek), with ALTO v2 XML, mean_ocr and per-word wc values — large enough for pretraining, domain adaptation, and systematic error analysis, but not a gold-standard transcription corpus.
  • IIIF-first images: images are the original IIIF-served scans (varying resolutions by holding library); the dataset includes page_iiif_url and image width/height so you can stream or fetch only what you need.
  • Sampling + metadata workflow: a small metadata config lets you pick pages before downloading large rows; many workflows stream images and avoid loading everything into RAM.
Who it's for

Great fit if you want realistic, historical OCR training or evaluation data (line-level supervision, confidence-aware filtering, layout and reading-order models) or if you need paired image+ALTO to compare modern OCR/VLM outputs against legacy OCR. Look elsewhere if you need gold-standard transcriptions for evaluation, a fully representative sample of Europeana, or consistently high-resolution images — OCR quality and image sizes vary by collection, and some pages have alignment issues marked in box_alignment.

Where it fits

Use this as silver-scale data for pretraining, error-finding and re-OCR pipelines, or as a diverse visual+text sample for document-understanding model development; complement with smaller hand-corrected datasets for final benchmarks.

Information

  • Websitehuggingface.co
  • OrganizationsBigLAM initiative, Hugging Face, Europeana, Austrian National Library, National Library of Finland, Hamburg State Library, University of Belgrade, Berlin State Library, National Library of Estonia, National Library of Poland, National Library of Luxembourg
  • AuthorsDaniel van Strien
  • Published date2026/09/30

Categories

More Items

Hugging Face

De-identified longitudinal multimodal CT dataset for multicancer screening that pairs ~24k CT volumes with radiology reports and voxel-wise tumor annotations across 13 cancer types. Designed for longitudinal disease modeling, detection/segmentation and vision–language research; CC BY‑NC‑ND 4.0 for non-commercial use.

Hugging Face

Provides 1.21M densely annotated desktop screenshots and 159.7M element instances for training and evaluating GUI grounding and screen-parsing models. Includes per-element accessibility-derived annotations, 917K recorded click transitions, multi-application scenes across seven appearance presets and resolutions; distributed as WebDataset shards with Parquet indexes.

Hugging Face

Generates humanoid robot motion references that preserve object contact locations/timing by solving windowed trajectory optimizations against contact targets in the object frame. Releases retargeted trajectories for two Unitree robots across 75 objects (≈13.9k robot–motion pairs); CC BY‑NC‑SA 4.0.