paper web signal

UndoBench: agents complete 83.54% of tasks, recover 46.72%

TL;DR

  • UndoBench reports tool-using AI agents reach 83.54% nominal task completion but only 46.72% conditional recovery success on enterprise workflow faults.
  • Naive retry produces duplicate external effects in 53.33% of trials, meaning simple retry logic can double user-visible side effects after a fault.
  • The benchmark covers 5,760 executions and 2,880 paired trials across 36 workflows, 36 fault scenarios, 8 enterprise domains, two open-weight models and two frameworks.

A new benchmark reports tool-using AI agents complete 83.54% of enterprise workflows but recover from only 46.72% of mid-execution faults, with naive retry producing duplicate external effects in 53.33% of trials.

The paper, UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents, by Dolly Sah, Tanmay Sah, Harshul Jain and Tanya Sah, argues standard evaluation misses the operational half of agent behavior. "Tool-using AI agents are increasingly deployed across enterprise software systems, yet widely used benchmarks primarily evaluate nominal task completion, conflating baseline planning competence with operational fault recovery," the abstract states.

The run is large. The authors report 5,760 executions and 2,880 paired trials, built from 36 base workflows and 36 fault scenarios across 8 enterprise domains, measured on 12 held-out test workflows across two open-weight models, two frameworks, and three recovery paradigms. Extensions to commercial API models reproduced the same competence-recovery separation, according to the abstract.

Recovery is phase-dependent, by the paper's reading. Before a mutation, methods perform similarly without duplicate effects among capable trials. During partial mutation, naive retry, per-call idempotency and zero-privilege journaling all collapse on the evaluated composite workflows. After commit but before acknowledgment, verification and server-side idempotency substantially improve safety. The authors conclude that "evaluating nominal completion alone masks critical, phase-dependent recovery vulnerabilities" in autonomous agents.

The abstract does not name the specific open-weight models, frameworks or commercial APIs tested.