Artificial Analysis Ships Index v4.3, Adds AutomationBench-AA
TL;DR
- Intelligence Index v4.3 replaces τ³-Banking with a 657-task AutomationBench-AA built with Zapier and upgrades Terminal-Bench to v4.0.
- Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) tie at 53, with Claude Opus 5 at 51 and Muse Spark 1.3 at 48.
- GPT-6 Astra averages $3.26 per Index task versus Fable's $7.63, a 57% lower cost for the same top-line score.
The new Intelligence Index v4.3, published September 7, 2026, replaces τ³-Banking with AutomationBench-AA and upgrades Terminal-Bench from v2.1 to v4.0. Private test set evaluations now count for 45% of the composite, up from 40%.
"In collaboration with Zapier, we run the held-out test set of 657 tasks, using the v1.0.6 version of the benchmark," the article says of AutomationBench-AA. "Its 657 tasks span Finance, HR, Marketing, Operations, Sales, and Support."
At the top of the leaderboard, Claude Fable 5.1 (max with fallback) and GPT-6 Astra (max) tie at 53. Claude Opus 5 (max) sits at 51 and Muse Spark 1.3 (max) at 48. The gap between the two leaders shows up in the bill. GPT-6 Astra averages $3.26 per Intelligence Index task against Fable's $7.63, a difference the article calls "57% lower for Astra." On AutomationBench-AA specifically, Astra scored 68.5% on objectives and cleared 41.6% of workflows without a guardrail violation.
The scoring rules matter. AutomationBench-AA counts a score as the "share of task objectives a model completes, where any task with a guardrail violation scores zero." As the piece puts it, "Completing every objective while respecting all guardrails remains harder than completing part of a workflow."
Artificial Analysis frames v4.3 as an appetizer for the next major version: "We are continuing to prioritize keeping Intelligence Index as useful as possible by bringing forward a subset of the changes we had planned for Index v5." Terminal-Bench v4.0, per the same article, "recalibrates compute and time allowances and improves task instructions, environments, and verification."
The article never says why τ³-Banking was retired, only that AutomationBench-AA moves to broader business workflows. It is one more entry on our agents beat, where we have logged 410 stories over the past 90 days.
Originally reported by artificialanalysis.ai
Read the original article →Original headline: Artificial Analysis Ships Intelligence Index v4.3 With AutomationBench-AA and Terminal-Bench 4.0