artificialanalysis.ai web signal

Claude Sonnet 5.5 Ranks #2 at 56, With Record Token Burn

Anthropic ai-business

TL;DR

  • Claude Sonnet 5.5 scored 56 on Artificial Analysis's Intelligence Index, ranking #2, two points behind Opus 5.5 (max).
  • At max effort it used ~193k output tokens per benchmark task, roughly 7x GPT-6 Astra and the highest the tracker has ever measured.
  • Per-token pricing is unchanged from Sonnet 5 at $2/$10 per million input/output, but cost per Intelligence Index task rose to $7.60.

Anthropic's Claude Sonnet 5.5 scored 56 on Artificial Analysis's Intelligence Index, landing at #2, two points behind Opus 5.5 (max). To get there, it used more output tokens per task than any model the tracker has measured.

"At max effort, where it reaches performance nearing that of Opus 5.5, Claude Sonnet 5.5 used ~193k Output Tokens per Intelligence Index Task," Artificial Analysis wrote. "This is the highest token use we have measured, around 60% higher than Opus 5.5 (max) or Sonnet 5 (max) and ~7x GPT-6 Astra (max)."

Per-token pricing is unchanged from Sonnet 5 at $0.2/$2/$10 per million cache input, input and output tokens. The bill still moves. Cost per Intelligence Index task lands at $7.60, roughly 50% higher than Sonnet 5's.

Sonnet 5.5 leads on coding. It reaches 64% on Terminal-Bench 4.0 against 60% for both Opus 5.5 and GPT-6 Astra. It lags Opus 5.5 on factual knowledge, scoring 54% on AA-Omniscience against 66%, with a hallucination rate of 47% versus 59%.

The token math tracks with a warning Anthropic itself surfaced this week when it published an Opus 5.5 prompting guide, cautioning that old Opus 5 effort settings now blow out tokens on the newer models.

Shared on Bluesky by 1 AI expert