artificialanalysis.ai via Hacker News

Kimi K3 Takes Second on AA-Briefcase, Behind Claude Fable 5

TL;DR

  • Moonshot's Kimi K3, released July 17 2026, scores 1543 Elo on AA-Briefcase, second only to Claude Fable 5's 1574.
  • The 2.8T-parameter model passes 51% of rubric checks and posts a 1754 analytical-quality Elo, edging Fable 5's 1744.
  • Each task costs $10.57 and takes 56.4 minutes on average, roughly 2.5x slower than Fable 5.

Moonshot's Kimi K3 lands in second place on Artificial Analysis's AA-Briefcase leaderboard, and the interesting part isn't the ranking, it's how close an open-weight model has gotten to the frontier. K3 posts an Elo of 1543 against Claude Fable 5's 1574, and on analytical quality specifically it actually edges Fable 5, 1754 to 1744.

The jump from the prior generation is the number that made me sit up. Artificial Analysis puts K3's improvement over Kimi K2.6 at +727 Elo points, from 816 to 1543. That is not a normal generational bump for a benchmark that's been running long enough to stabilize. It puts K3 clear of GPT-5.6 Sol at 1501, Claude Sonnet 5 at 1388, and Claude Opus 4.8 at 1347, so on this particular test the open-weight camp has essentially caught the closed labs on everything except the very top slot.

Where K3 pays for that is the runtime and the bill. Each task costs $10.57 on average and takes 56.4 minutes to complete, roughly 2.5x Fable 5's wall clock. K3 also chews through more of the task by volume, 83 turns and 120k output tokens per task versus 42k for K2.6. Presentation is the other soft spot, an Elo of 1471 that trails GPT-5.6 Sol's 1660 by a wide margin and even sits below Opus 4.8's 1492.

The honest caveat is that a benchmark is one lens. Artificial Analysis doesn't tell us what hardware Moonshot expects buyers to run a 2.8T model on, whether the extra wall-clock is a deliberate reasoning-depth choice or a training artifact, or how the presentation gap shows up in real customer-facing work. The $3 in and $15 out per million tokens, with a 90% cached-token discount, hints that repeated agent flows could look better than the headline $10.57 suggests, but that's the buyer's calculation to run.

What this changes for anyone shopping agents is the shape of the trade. Analytical depth without vendor lock-in is now on the table from Moonshot at roughly parity with Anthropic on the quality axis that matters most for research-style workloads. The bet you are making is that the cost and latency premium is worth the openness, and that Moonshot closes the presentation gap in the next revision.