huggingface.co web signal

SafeActBench Paper Shows Tool-Using Agents Pass 97% Static but Act Before Evidence 37-67% of Time

Agents Safety OpenAI ai-business

Summary

The paper 'From Evidence to Action' releases SafeActBench — 656 cases across six domains and five protocols — and finds frontier agents can judge actions correctly on static tests (GLM-ZCode 97.7%, DeepSeek-DSH 96.5%) yet act before evidence 37-67% of the time in interactive execution. Single actions succeed over 93% of the time once evidence is complete; multi-step DAG workflows crater (GLM 12.1%), with harness choice flipping failure modes rather than fixing them.