Expert attention map

The Who's Who of AI

What credible people across AI noticed, why it matters, and where the field is converging or disagreeing.

2,364 searchable experts 2,966 tracked across all sources
Clear
Showing signals surfaced by Daniel van Strien ×

The developments commanding sustained expert attention

One card per development. Sources are clustered; expert reactions remain attributable.

Established AI field signal Resource 8d ago
HuggingFace uv-scripts OCR tool

uv-scripts/ocr · Datasets at Hugging Face

1 directory member surfaced this signal.

1 expert 1 community 1 sources clustered

“Small open models are getting genuinely good at document parsing: OvisOCR2 (0.9B, Apache 2.0) is claiming SOTA on OmniDocBench v1.6. Day-1 recipe: OCR a whole HF image dataset — digitised newspapers, archives, zines — to markdown with one command. huggingfa…”

Established AI field signal Resource 14d ago
Apollo 11 audio transcription dataset

Apollo 11 — Search the Mission Audio - a Hugging Face Space by davanstrien

1 directory member surfaced this signal.

1 expert 1 community 1 sources clustered

“It's unedited machine output over scratchy radio (it keeps hearing "Follow eleven, this is Houston"). The recipe runs on any audio collection in one command: huggingface.co/datasets/uv-...”

Established Evaluation & benchmarks Analysis 16d ago
OCR benchmarking challenges essay

Benchmarking OCR fairly is harder than it looks – Daniel van Strien

2 experts are actively discussing the implications.

1 expert 1 community 1 sources clustered

“Write-up + plus how I ran the whole thing on HF Jobs, no local GPU: danielvanstrien.xyz/posts/2026/o...”

2 experts discussed this · 7 posts
Daniel van Strien: I ran 10 newer OCR models on @ai2.bsky.social's olmOCR-bench "old scans" subset. The ranking flips depending on what you actually want.
Daniel van Strien: On the headline score, PaddleOCR-VL beats NuExtract3 (38.6 vs 37.8). But rank by how much of the page each model actually reads, and NuExtract3 is well ahead (41.6 vs 31.2). Same two models, opposi…
Daniel van Strien: The score rewards dropping boilerplate, i.e. letterheads, stamps, page numbers, so a model that reads the page more faithfully can rank lower.
Open the full discussion →