artificialanalysis.ai web signal

Artificial Analysis Ships Index v4.3, Adds AutomationBench-AA

OpenAI Anthropic Agents ai-business

TL;DR

  • Intelligence Index v4.3 replaces τ³-Banking with a 657-task AutomationBench-AA built with Zapier and upgrades Terminal-Bench to v4.0.
  • Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) tie at 53, with Claude Opus 5 at 51 and Muse Spark 1.3 at 48.
  • GPT-6 Astra averages $3.26 per Index task versus Fable's $7.63, a 57% lower cost for the same top-line score.

The new Intelligence Index v4.3, published September 7, 2026, replaces τ³-Banking with AutomationBench-AA and upgrades Terminal-Bench from v2.1 to v4.0. Private test set evaluations now count for 45% of the composite, up from 40%.

"In collaboration with Zapier, we run the held-out test set of 657 tasks, using the v1.0.6 version of the benchmark," the article says of AutomationBench-AA. "Its 657 tasks span Finance, HR, Marketing, Operations, Sales, and Support."

At the top of the leaderboard, Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) tie at 53. Claude Opus 5 (max) sits at 51 and Muse Spark 1.3 (max) at 48. The gap between the two leaders shows up in the bill. GPT-6 Astra averages $3.26 per Intelligence Index task against Fable's $7.63, a difference the article calls "57% lower for Astra." On AutomationBench-AA specifically, Astra scored 68.5% on objectives and cleared 41.6% of workflows without a guardrail violation.

The scoring rules matter. AutomationBench-AA counts a score as the "share of task objectives a model completes, where any task with a guardrail violation scores zero." As the piece puts it, "Completing every objective while respecting all guardrails remains harder than completing part of a workflow."

Artificial Analysis frames v4.3 as an appetizer for the next major version: "We are continuing to prioritize keeping Intelligence Index as useful as possible by bringing forward a subset of the changes we had planned for Index v5." Terminal-Bench v4.0, per the same article, "recalibrates compute and time allowances and improves task instructions, environments, and verification."

The article never says why τ³-Banking was retired, only that AutomationBench-AA moves to broader business workflows. It is one more entry on our agents beat, where we have logged 410 stories over the past 90 days.