OCR benchmarking challenges essay
Benchmarking OCR fairly is harder than it looks – Daniel van Strien
2 experts are actively discussing the implications.
1 expert
1 community
1 sources clustered
“Write-up + plus how I ran the whole thing on HF Jobs, no local GPU: danielvanstrien.xyz/posts/2026/o...”
2 experts discussed this · 7 posts
Daniel van Strien: I ran 10 newer OCR models on @ai2.bsky.social's olmOCR-bench "old scans" subset. The ranking flips depending on what you actually want.
Daniel van Strien: On the headline score, PaddleOCR-VL beats NuExtract3 (38.6 vs 37.8). But rank by how much of the page each model actually reads, and NuExtract3 is well ahead (41.6 vs 31.2). Same two models, opposi…
Daniel van Strien: The score rewards dropping boilerplate, i.e. letterheads, stamps, page numbers, so a model that reads the page more faithfully can rank lower.
Open the full discussion →