Found first: a primary source the press has not covered yet.
A new benchmark called UndoBench finds enterprise AI agents complete 83.54% of tasks under normal conditions, but recover correctly from mid-execution faults in only 46.72% of cases. Naive retry compounds the problem, producing duplicate external effects in 53.33% of trials. The paper is at arxiv.org/abs/2610.05622.
What the source says
Independent researchers Dolly Sah, Tanmay Sah, Harshul Jain, and Tanya Sah built UndoBench from 36 base workflows and 36 fault scenarios across 8 enterprise domains, generating 5,760 total executions across 2,880 counterfactual paired trials. Each trial pair uses identical seeds, isolating recovery capability from task competence. They tested two open-weight models across two agent frameworks and three recovery paradigms, reporting conditional recovery success rate (CRSR) as the primary recovery metric. Extensions to commercial API models reproduced the same competence-recovery gap.
Why it matters
Standard benchmarks measure whether an agent finishes a task. UndoBench measures what happens when something fails partway through, treating these as separate skills. An agent that scores 83.54% on task completion recovers correctly from mid-execution faults in fewer than half of cases. The naive retry finding has direct operational weight: retrying a partially-executed workflow produces duplicate external effects in more than half of trials, turning a recoverable failure into a compounding one. The gap held across all tested combinations of model, framework, and recovery paradigm.