artificialanalysis.ai via Hacker News

Artificial Analysis v4.2 Doubles Private Test Weight to 40%

Anthropic OpenAI ai-business

TL;DR

  • Private held-out test sets now carry 40% weight in the v4.2 Intelligence Index, double their share in v4.1.
  • New GDP.pdf eval, built with Surge AI, spans 4,592 PDF pages graded against 1,275 expert-authored criteria.
  • Claude Fable 5.1 leads the refreshed leaderboard, GPT-6 Astra is second, Meta third; saturated GPQA Diamond is retired.

Private, held-out test sets now count for 40% of the Artificial Analysis Intelligence Index, double their v4.1 weight. That is the headline change in v4.2, published September 4 by Artificial Analysis, and the firm's stated reason for it is blunt: the shift is meant to "reduce the ability for labs to game evaluations."

Two new evaluations ride that expansion in. AA-Briefcase is a private agentic suite that, per Artificial Analysis, "tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files." The second, GDP.pdf, is built with Surge AI and "evaluates single-turn professional document reasoning across 100 PDFs and ten domains. Models must synthesize evidence distributed across 4,592 pages, including text, tables, charts, footnotes, and exclusions." Answers are graded against 1,275 expert-authored criteria.

GPQA Diamond is retired. The firm calls it saturated.

The refreshed leaderboard puts Anthropic's Claude Fable 5.1 first overall, with OpenAI's GPT-6 Astra second on a roughly 4-point gain over GPT-5.6 Sol. Meta lands third, followed by SpaceXAI, Moonshot/Kimi, Z.AI and Google. On the individual GDP.pdf slice, the top order flips: GPT-6 Astra scores 33.2%, GPT-5.6 Sol 28.2%, Claude Fable 5.1 26.2%. On AA-Briefcase, GPT-6 Astra shows roughly an 85 Elo-point lead over GPT-5.6 Sol.

Artificial Analysis frames the release as a stopgap, writing that it is "accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier." The timing is not neutral for either lab at the top of the board: Anthropic's IPO marketing has slipped to mid-October, and GPT-6 Astra is one of a run of OpenAI stories we have tracked this week across our OpenAI feed.