arxiv.org web signal

Argo-Bench: Top Data Agent Clears 95+ on Just 34.8% of Tasks

Enterprise AI Agents ai-business

TL;DR

  • Argo-Bench grades 210 analytics tasks on a simulated New York food-delivery warehouse of 235 Oracle E-Business Suite tables and 7.5 billion rows.
  • The strongest of 14 frontier and open-weight models scored 95 or higher on only 34.8% of tasks and averaged 59.5 points.
  • The grader scores consequences inside the simulator, such as fraud losses prevented and forecasts against held-out months, not query correctness.

The strongest of 14 frontier and open-weight models cleared 95 or above on only 34.8% of the 210 tasks in Argo-Bench, a benchmark posted to arXiv on October 1 by TextQL Labs that stress-tests data-science agents against a simulated 7.5-billion-row enterprise warehouse. Average score across the field was 59.5, with Claude Opus 5.5 at the top.

Argo-Bench models "a food delivery platform in New York City in 2024" with 81 million orders and exports it to "an Oracle E-Business Suite warehouse of 235 tables and 7.5 billion rows." What separates it from text-to-SQL benchmarks is where the credit lives: the grader scores "the consequences of an action rather than the correctness of a query, using the simulator's latent state as ground truth." Fraud-ring bans are rated by the losses they prevent. Forecasts are tested against held-out months.

The paper catalogues four categories of failure. Agents picked the wrong record (missing that "a member whose card is declined keeps benefits for seven days"), the wrong objective (cutting "where quests buy the fewest courier-hours" rather than where they "save the least surge per bonus dollar"), and the wrong quantity (46 of them missed "all 12 monthly counts" of location check-ins by polling hourly while the app sends half-hourly). Forecast calibration was worst of all: across 4,553 series, "nominal 80% intervals contain the realized value only 44.8% of the time."

Argo-Bench lands into an evaluation space our tracker has been logging heavily, with 424 agent stories in the past 90 days, but is among the few to grade on business consequences rather than output text. The harness, reference agents and grader are on GitHub under Apache 2.0; the warehouse is on Hugging Face under CC BY 4.0.