Daniel van Strien
Machine Learning Librarian at Hugging Face
Machine Learning Librarian at Hugging Face with public evidence across AI research.
- AI signals
- 9 past 30d
- Sources
- 3 distinct domains
- Discussões
- 1 past 30d
- Latest signal
- 4d ago
Articles & links
Need a historical illustration? Search 1.49 million images from British Library books and Britannica (1500s–1920s). Find similar images, download full-resolution images or transparent PNG cutouts, and explore the sources. huggingface.co/spaces/davan...
Got a digitised collection that needs OCR? uv-scripts is a set of single-file Python scripts that OCR a whole image dataset to markdown in one command — 20+ open VLMs to pick from, nothing to install but uv. github.com/davanstrien/...
- Each script is a self-contained Python file using PEP 723 inline dependency declarations, runnable with a single `uv run` command.
- Nine task categories are covered including OCR with 30+ models, audio transcription, vision detection, embeddings, and LLM inference.
- Scripts use standardized argument patterns so both humans and AI agents can run them locally or on Hugging Face Jobs GPU infrastructure.
The Europeana Newspapers dataset on @hf.co now has an `alto` config: the raw ALTO XML for all 5.9M pages. The coordinates for every word, line and block, per-word OCR confidence and font info that the flattened text dropped are back. huggingface.co/datasets/biglam/europeana_ne…
- The archive holds 17,831,085 newspaper page rows and roughly 32 billion tokens across 477 GB of parquet, curated by the BigLAM initiative.
- German dominates with 4.6 million rows; French, Estonian, Polish, Finnish and Swedish follow, while Yiddish (70) and Ukrainian (40) sit near the bottom.
- Each row includes mean_ocr and std_ocr confidence scores plus an IIIF link; a separate config exposes the raw ~420 GB ALTO v2 XML.
A new @hf.co and @eleutherai.bsky.social project, with @storytracer.com! Blog post: huggingface.co/blog/fineboo...
- Hugging Face and EleutherAI's FineBooks leaderboard scores 14 open OCR models against 2,165 expert-transcribed pages from the Biodiversity Heritage Library.
- dots.mocr (3B) tops the leaderboard at 97.6% reading accuracy, with sub-2B models OvisOCR2 and PaddleOCR-VL-1.6 close behind at under half a dollar per thousand pages.
- Leading models are good enough for LLM training corpora but silently modernize archaic letterforms, ruling them out for faithful scholarly transcription.
A 15.9M-parameter specialist OCR model outperforms many VLMs hundreds of times larger. On a 2,165-page historical eval dataset, Kraken PP-OCRv6 ranks #4 on reading CER and #1 when long-s, ligatures and case are preserved. huggingface.co/spaces/fineb...
- rednote-hilab's dots.mocr (3B) leads the FineBooks BHL leaderboard at 97.6% reading accuracy across 2,165 expert-transcribed pages.
- Sub-2B models OvisOCR2 (0.9B, 96.9%) and PaddleOCR-VL-1.6 (1B, 96.1%) come next at under $0.50 per thousand pages.
- Ground truth covers only Antiqua typefaces in English, French, German and Latin; Fraktur, handwriting and multi-column layouts are excluded.
OCR for Japanese manga, Swedish handwriting or Arabic print? There’s a growing range of OCR models on @hf.co: VLMs, dedicated text recognisers and complete OCR pipelines. I’ve gathered 41 models into four collections, with short notes to help you choose: huggingface.co/collect…
- davanstrien's OCR on the Hub groups models into four sub-collections covering documents, languages and scripts, handwriting and archives, and text recognition pipelines.
- The languages sub-collection covers Thai, Japanese, Vietnamese, Arabic, Korean and Devanagari; handwriting spans Swedish, Norwegian, German Kurrent, Tibetan, and Hebrew-script manuscripts.
- The text recognition sub-collection features Kraken and PaddleOCR pipelines for documents, manga and text in photographs.
Public data on which coding agents actually use the Hugging Face Hub, refreshed monthly. Claude Code now sends a majority of everything the Hub can attribute to a named agent. In May, it led on 4 days out of 30. In July, all 30. huggingface.co/datasets/hug...
- Claude Code led July 2026 Hub agent traffic at 44.4% of requests and 39.5% of distinct users, with Codex second at 20.8% and 21.2%.
- Hugging Face publishes only relative shares, not raw counts, so falling percentages do not indicate declining usage.
- Attribution runs through the huggingface_hub Python library's User-Agent token, so direct HTTP calls and unregistered harnesses are excluded.
Derived datasets are bigger on Hugging Face Hub than people realise. ~73% of analysed datasets on the Hub are derivatives of something else, i.e. cleaned, translated, extended, etc. Built an explorer that infers the missing lineage from content: huggingface.co/spaces/davan...
Uploaded a dataset of 1,080,814 public domain images, mostly from 19th-century books, to the Hugging Face Hub: huggingface.co/datasets/big... You can also do semantic search against the images here: huggingface.co/spaces/davan...
Uploaded a dataset of 1,080,814 public domain images, mostly from 19th-century books, to the Hugging Face Hub: huggingface.co/datasets/big... You can also do semantic search against the images here: huggingface.co/spaces/davan...
You can now run SQL over 2.19 BILLION web pages — zero download. @commoncrawl.bsky.social April 2026 crawl + URL index are on Hugging Face Storage Buckets. DuckDB reads it straight over hf:// — I counted all 2.19B in ~35s. Or point your own agent at it 👇 huggingface.co/spaces/…
Recent commentary
I think VLM-based OCR might finally be close to working on historic newspapers! Many models I've tried before failed i.e. hallucinations, repetition loops, context overflow. Surya OCR 2 (a 650M model!) does a very good job!
NuExtract3 (4B, Apache-2.0) does OCR *and* structured extraction. Point it at a dataset of scanned index cards + a JSON schema → clean catalog JSON One command on huggingface Jobs. (or skip the schema for plain Markdown OCR) Script + dataset 👇
If libraries, archives and museums pooled their (labelled) data, they could build state-of-the-art open models for the things they actually care about! I tried a small version: one open model (NuExtract-3, 4B) fine-tuned to read archival index cards across several collections.
FineBooks wants to help re-OCR the public domain to unlock a rich source of data for AI, for researchers, and for enjoyment. Step one: work out which modern OCR models are actually good enough.
What could a rich ecosystem of small GLAM AI models enable? IMO: cheaper, better-fitted, more robust models. Example: I extended an existing @natlibscot.bsky.social archival card detector to 4 collections to make a more generic index card detector. Took an hour or two and minimal $
In Daniel van Strien's orbit
Center = Daniel van Strien. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.
Are you Daniel van Strien? Show it.
Add the Who’s Who of AI badge to your site or bio. It links back to this profile.
Markdown: [](https://aiweekly.co/whos-who/person/daniel-van-strien)