sakana.ai web signal

Sakana AI ships Fugu Max, cheaper routing over open models

TL;DR

  • Fugu Max prices at $2 per million input tokens and $6 per million output tokens, which Sakana claims runs 40-60% below Sonnet 5, GPT 5.6 Terra, and Kimi K3.
  • The system routes each task across a swappable pool of open-weight and specialized models, with NVIDIA's Nemotron folded in via an August 2026 collaboration.
  • Fugu Ultra v2 scores 48.3 on Chartography against Opus 5's 27.3, and does so without Fable 5, Fable 5.1, or GPT-6-Astra in its agent pool.

Sakana AI shipped Fugu Max on September 11, an orchestration engine priced at $2 per million input tokens and $6 per million output tokens that routes each request to what the company calls "the leanest model capable of solving them." The pricing runs 40 to 60 percent below the output cost of Sonnet 5, GPT 5.6 Terra, and Kimi K3, according to the company's own comparison. Two experts in our Who's Who directory shared the release on launch day.

Sakana paired the launch with Fugu Ultra v2, a higher-capability variant built on the same routing architecture, and frames the two as questions with the same engine underneath. "Fugu Max asks: What is the best possible output we can deliver at the lowest possible cost?" the release states, while "Fugu Ultra v2 asks: What is the absolute highest capability we can achieve on complex, multi-step tasks?"

On self-reported benchmarks Fugu Max claims the best overall score on six evaluations including Terminal Bench 2.1, GPQAD, AA-LCR, GDP.pdf, AutomationBench and SWEFish. Ultra v2 posts 48.3 on Chartography against Opus 5's 27.3, and 74.3 on DeepSWE. No third-party evaluation is cited in the announcement.

The strategic pitch sits underneath the numbers. "By orchestrating a swappable pool of open and specialized models, it outperforms closed ecosystems while protecting users from vendor lock-in," the company writes, describing the design as "supply chain resilience by design" against API revocations, geopolitical turbulence, and sudden service cutoffs. NVIDIA's Nemotron models were folded into the pool through a collaboration Sakana announced in August. Ultra v2 "achieves these scores without Fable 5, Fable 5.1, or GPT-6-Astra in its agent pool," a line that reads as both a boast and a map of which frontier stacks Sakana can and cannot draw from.

Shared on Bluesky by 2 AI experts