Daniel van Strien

Machine Learning Librarian at Hugging Face

Why they matter

Machine Learning Librarian at Hugging Face with public evidence across AI research.

AI signals
9
past 30d
Sources
3
distinct domains
Discusiones
0
past 30d
Latest signal
5h ago
View every signal from Daniel van Strien →
Machine Learning Librarian at @hf.co

Articles & links

Got a digitised collection that needs OCR? uv-scripts is a set of single-file Python scripts that OCR a whole image dataset to markdown in one command — 20+ open VLMs to pick from, nothing to install but uv. github.com/davanstrien/...

GitHub - davanstrien/uv-scripts-for-ai: Self-contained UV scripts for data & ML tasks — OCR, vision, audio & more — run one in a command, locally or on Hugging Face Jobs. Built for humans and agents. github.com
AI Weekly's analysis
  • Each script is a self-contained Python file using PEP 723 inline dependency declarations, runnable with a single `uv run` command.
  • Nine task categories are covered including OCR with 30+ models, audio transcription, vision detection, embeddings, and LLM inference.
  • Scripts use standardized argument patterns so both humans and AI agents can run them locally or on Hugging Face Jobs GPU infrastructure.
Read full analysis →
View on Bluesky · ♥ 38 ↻ 13 ↩ 1 · 2 from the directory shared this · 89d ago

A new @hf.co and @eleutherai.bsky.social project, with @storytracer.com! Blog post: huggingface.co/blog/fineboo...

FineBooks: are open OCR models good enough to unlock historical knowledge? huggingface.co
AI Weekly's analysis
  • Hugging Face and EleutherAI's FineBooks leaderboard scores 14 open OCR models against 2,165 expert-transcribed pages from the Biodiversity Heritage Library.
  • dots.mocr (3B) tops the leaderboard at 97.6% reading accuracy, with sub-2B models OvisOCR2 and PaddleOCR-VL-1.6 close behind at under half a dollar per thousand pages.
  • Leading models are good enough for LLM training corpora but silently modernize archaic letterforms, ruling them out for faithful scholarly transcription.
Read full analysis →
View on Bluesky · ♥ 4 ↻ 2 ↩ 1 · 2 from the directory shared this · 28d ago

A 15.9M-parameter specialist OCR model outperforms many VLMs hundreds of times larger. On a 2,165-page historical eval dataset, Kraken PP-OCRv6 ranks #4 on reading CER and #1 when long-s, ligatures and case are preserved. huggingface.co/spaces/fineb...

BHL OCR Leaderboard - a Hugging Face Space by finebooks huggingface.co
AI Weekly's analysis
  • rednote-hilab's dots.mocr (3B) leads the FineBooks BHL leaderboard at 97.6% reading accuracy across 2,165 expert-transcribed pages.
  • Sub-2B models OvisOCR2 (0.9B, 96.9%) and PaddleOCR-VL-1.6 (1B, 96.1%) come next at under $0.50 per thousand pages.
  • Ground truth covers only Antiqua typefaces in English, French, German and Latin; Fraktur, handwriting and multi-column layouts are excluded.
Read full analysis →
View on Bluesky · ♥ 32 ↻ 6 ↩ 2 · 2 from the directory shared this · 4d ago

OCR for Japanese manga, Swedish handwriting or Arabic print? There’s a growing range of OCR models on @hf.co: VLMs, dedicated text recognisers and complete OCR pipelines. I’ve gathered 41 models into four collections, with short notes to help you choose: huggingface.co/collect…

OCR on the Hub - a davanstrien Collection huggingface.co
AI Weekly's analysis
  • davanstrien's OCR on the Hub groups models into four sub-collections covering documents, languages and scripts, handwriting and archives, and text recognition pipelines.
  • The languages sub-collection covers Thai, Japanese, Vietnamese, Arabic, Korean and Devanagari; handwriting spans Swedish, Norwegian, German Kurrent, Tibetan, and Hebrew-script manuscripts.
  • The text recognition sub-collection features Kraken and PaddleOCR pipelines for documents, manga and text in photographs.
Read full analysis →
View on Bluesky · ♥ 19 ↻ 7 ↩ 0 · 2 from the directory shared this · 5h ago

Public data on which coding agents actually use the Hugging Face Hub, refreshed monthly. Claude Code now sends a majority of everything the Hub can attribute to a named agent. In May, it led on 4 days out of 30. In July, all 30. huggingface.co/datasets/hug...

huggingface/agent-usage · Datasets at Hugging Face huggingface.co
AI Weekly's analysis
  • Claude Code led July 2026 Hub agent traffic at 44.4% of requests and 39.5% of distinct users, with Codex second at 20.8% and 21.2%.
  • Hugging Face publishes only relative shares, not raw counts, so falling percentages do not indicate declining usage.
  • Attribution runs through the huggingface_hub Python library's User-Agent token, so direct HTTP calls and unregistered harnesses are excluded.
Read full analysis →
View on Bluesky · ♥ 13 ↻ 2 ↩ 4 · 2 from the directory shared this · 35d ago

Small open models are getting genuinely good at document parsing: OvisOCR2 (0.9B, Apache 2.0) is claiming SOTA on OmniDocBench v1.6. Day-1 recipe: OCR a whole HF image dataset — digitised newspapers, archives, zines — to markdown with one command. huggingface.co/datasets/uv-sc…

uv-scripts/ocr · Datasets at Hugging Face huggingface.co
View on Bluesky · ♥ 34 ↻ 2 ↩ 0 · 55d ago

Recent commentary

I think VLM-based OCR might finally be close to working on historic newspapers! Many models I've tried before failed i.e. hallucinations, repetition loops, context overflow. Surya OCR 2 (a 650M model!) does a very good job!

View on Bluesky · ♥ 87 ↻ 14 ↩ 1 · 77d ago

NuExtract3 (4B, Apache-2.0) does OCR *and* structured extraction. Point it at a dataset of scanned index cards + a JSON schema → clean catalog JSON One command on huggingface Jobs. (or skip the schema for plain Markdown OCR) Script + dataset 👇

View on Bluesky · ♥ 39 ↻ 13 ↩ 2 · 110d ago

If libraries, archives and museums pooled their (labelled) data, they could build state-of-the-art open models for the things they actually care about! I tried a small version: one open model (NuExtract-3, 4B) fine-tuned to read archival index cards across several collections.

View on Bluesky · ♥ 28 ↻ 5 ↩ 1 · 75d ago

FineBooks wants to help re-OCR the public domain to unlock a rich source of data for AI, for researchers, and for enjoyment. Step one: work out which modern OCR models are actually good enough.

View on Bluesky · ♥ 25 ↻ 5 ↩ 2 · 28d ago

What could a rich ecosystem of small GLAM AI models enable? IMO: cheaper, better-fitted, more robust models. Example: I extended an existing @natlibscot.bsky.social archival card detector to 4 collections to make a more generic index card detector. Took an hour or two and minimal $

View on Bluesky · ♥ 14 ↻ 1 ↩ 1 · 97d ago

In Daniel van Strien's orbit

Center = Daniel van Strien. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.

Are you Daniel van Strien? Show it.

Add the Who’s Who of AI badge to your site or bio. It links back to this profile.

Listed in AI Weekly's Who's Who of AI

Markdown: [![Listed in AI Weekly's Who's Who of AI](https://aiweekly.co/modules/custom/aiweekly_whoswho/images/whoswho-badge.svg)](https://aiweekly.co/whos-who/person/daniel-van-strien)