SafeActBench Paper Shows Tool-Using Agents Pass 97% Static but Act Before Evidence 37-67% of Time
Summary
The paper 'From Evidence to Action' releases SafeActBench — 656 cases across six domains and five protocols — and finds frontier agents can judge actions correctly on static tests (GLM-ZCode 97.7%, DeepSeek-DSH 96.5%) yet act before evidence 37-67% of the time in interactive execution. Single actions succeed over 93% of the time once evidence is complete; multi-step DAG workflows crater (GLM 12.1%), with harness choice flipping failure modes rather than fixing them.
Originally reported by huggingface.co
Read the original article →Original headline: SafeActBench Paper Shows Tool-Using Agents Pass 97% Static but Act Before Evidence 37-67% of Time