BigLAM packs 32B-token Europeana newspapers on Hugging Face
TL;DR
- The archive holds 17,831,085 newspaper page rows and roughly 32 billion tokens across 477 GB of parquet, curated by the BigLAM initiative.
- German dominates with 4.6 million rows; French, Estonian, Polish, Finnish and Swedish follow, while Yiddish (70) and Ukrainian (40) sit near the bottom.
- Each row includes mean_ocr and std_ocr confidence scores plus an IIIF link; a separate config exposes the raw ~420 GB ALTO v2 XML.
The Hugging Face dataset biglam/europeana_newspapers bundles historic European newspapers into a single machine-readable archive: 17,831,085 rows, 477 GB of parquet, and roughly 32 billion tokens.
The collection was assembled by the BigLAM initiative from Europeana's newspaper full-text dumps and spans, in the card's words, "the 18th to the early 20th century." German dominates the mix with 4.6 million rows. French follows at 107k, Estonian at 497k, Polish at 108k, Finnish at 85.4k and Swedish at 46.3k. Greek, Russian, Croatian, Serbian, Ukrainian and Yiddish appear in far smaller counts. Yiddish has 70 rows. Ukrainian has 40.
Every row carries the extracted text along with mean_ocr and std_ocr confidence scores, bounding boxes for illustrations, publication date, newspaper title, and an IIIF URL back to the original digitized image. A separate config exposes the raw ALTO v2 XML at roughly 420 GB, with per-word OCR confidence, coordinates and font information. Two of the AI researchers on our Who's Who watchlist have already posted the link.
The card is upfront about the limits. On OCR: "Errors present, especially in older newspapers or non-standard fonts." On content: "Reflects prejudices and perspectives of source time periods; may contain offensive content by modern standards." Coverage is temporally and linguistically uneven. The corpus traces back to Europeana's full-text dumps of March 2019, so any digitization done after that date is not here. Dataset-card contact is listed as Daniel van Strien at Hugging Face.
Shared on Bluesky by 2 AI experts
-
The Europeana Newspapers dataset on @hf.co now has an `alto` config: the raw ALTO XML for all 5.9M pages. The coordinates for every word, line and block, per-word OCR confidence and font info that the flattened text dro…
View on Bluesky →
Originally reported by huggingface.co
Read the original article →Original headline: biglam/europeana_newspapers · Datasets at Hugging Face