Hugging Face dataset bundles 98,877 Europeana newspaper scans
TL;DR
- The dataset pairs 98,877 newspaper pages from the 1700s–1940s with OCR text, ALTO XML layout, and per-word confidence scores, drawn from eight European libraries.
- German pages dominate at 34,052, followed by Estonian (12,296) and Serbian (11,416); the sample spans ten languages across Fraktur, Cyrillic, and Greek scripts.
- The dataset card warns the historical OCR is 'suitable for pre-training or finding hard pages, not as benchmark truth.'
98,877 historical newspaper page images are now up on Hugging Face, bundled with the OCR text, ALTO XML layout, and per-word confidence scores that libraries ran against them. The pages span the 1700s to the 1940s and come from eight European institutions, with German pages (34,052) heaviest, followed by Estonian (12,296) and Serbian (11,416). Daniel van Strien, Hugging Face's Machine Learning Librarian, sampled and repackaged the set from Europeana's 5.9 million-page archive. Two of the researchers we follow in our tracker shared the link.
The dataset card is unusually frank about what readers are getting. The OCR was "Produced during digitization by libraries using ABBYY FineReader Engine" (65,636 pages) or a CCS docWorks workflow (33,241 pages), and is "suitable for pre-training or finding hard pages, not as benchmark truth." Per-word confidence scores are "Not comparable across collections" because engine settings differed. Some Serbian Cyrillic pages were OCR'd as Latin, producing "unreadable text with plausible boxes."
A further 637 pages carry alignment warnings, 610 of them from the Austrian National Library, where the ALTO coordinates drift off the text they describe. The sample is non-representative by design: any single newspaper title is capped at 5% of pages, with one page per issue preferred. The repo is labelled a proof of concept at 271 GB.
Shared on Bluesky by 2 AI experts
-
Uploaded a dataset of 98,877 historical newspaper pages (1700s–1940s) to the Hub, each with its original OCR text, word boxes and confidence scores. huggingface.co/datasets/big...
View on Bluesky →
Originally reported by huggingface.co
Read the original article →Original headline: biglam/europeana_newspapers_images · Datasets at Hugging Face