CheckerBench: Top Agent Clears 45.33% of 300 Checker Tasks
TL;DR
- On 300 CheckerBench tasks drawn from 297 CVEs across 167 repositories, Claude Opus 4.8 with OpenCode leads at 45.33% Pass@1.
- Mean Pass@1 across 21 model-harness configurations is 32.30%; OpenCode scores 36.95% versus the Hermes Agent at 26.14%.
- C/C++ is the hardest ecosystem: DeepSeek-V4-Pro leads it at only 24.53%, while Claude Opus 4.8 reaches 65.36% on Python.
The best coding agent clears less than half of a new benchmark built to test whether language models can build static-analysis checkers from scratch. On 300 executable tasks drawn from 297 CVEs across 167 repositories, Claude Opus 4.8 paired with the OpenCode harness tops the field at 45.33% Pass@1. The mean across 21 model-harness configurations is 32.30%.
CheckerBench, introduced on Hugging Face by researchers at East China Normal University, Shanghai Jiao Tong University, Peking University and partners, covers 85 CWE categories across five language ecosystems (C/C++, Java, Python, JavaScript, Go) and spans both the Clang Static Analyzer and CodeQL. Each task hands the agent a vulnerable revision, a fixed revision, a pinned analysis environment, and a checker scaffold.
The abstract is deadpan: 'reliable, reusable checker development remains challenging for current coding agents.' More than half of the 300 tasks fail even at the top of the leaderboard.
Harness matters almost as much as model. Mean Pass@1 is 36.95% under OpenCode, 33.81% under Claude Code, and 26.14% under the Hermes Agent. By language, C/C++ is where things fall apart: the leader there, DeepSeek-V4-Pro, hits only 24.53%, while Python climbs to 65.36% for Claude Opus 4.8.
Where failures land is also stark. Across 4,360 failed runs the paper labels 25.32% as budget exhaustion, 22.96% as inadequate synthesis, 21.15% as inadequate semantic modeling, and 15.96% as excess false positives. CheckerBench joins a thick run of coding-agent benchmarks in our tracker this week, and it is the one that makes the gap between harness scaffolding and real agent competence hard to ignore.
Originally reported by huggingface.co
Read the original article →Original headline: CheckerBench: Best Coding Agent Clears Just 45.3% of 300 Real-World Static-Analysis Checker Tasks