Sakana rebuilds Fugu conductor on Gemma 4, matches Qwen build
TL;DR
- Sakana AI retrained its Fugu conductor model on Gemma 4 E2B and reports comparable performance to the earlier Qwen-based version.
- Fugu is a two-layer setup: a small conductor routes work to a pool of frontier models the company calls hundreds-of-billions-scale.
- Sakana frames the result as base-model-agnostic training, groundwork for a future conductor built on a domestically-developed base model for sovereignty use cases.
Sakana AI has retrained its Fugu "conductor" model on Gemma 4 E2B and, according to the company's own post, the new build lands in roughly the same place as the earlier version that used Qwen as its base. The framing is deliberately modest: not a leap in capability, but evidence that Sakana's orchestration training method transfers across model lineages.
Fugu is a two-layer system. A small conductor model, trained by Sakana, sits at a single endpoint and decides how to route work. Behind it is a pool of much larger frontier models, which the company calls "hundreds of billions-scale," that handle the actual generation. The pitch to buyers is that you swap out the routing brain rather than the compute, so cost and sovereignty concerns live in the small piece you actually control.
That is where the Gemma 4 result matters. Sakana's earlier conductor was Qwen-based; the new one uses Gemma 4 E2B under Apache 2.0. Being able to rebuild on a different base lineage and reach "comparable performance and equivalent cost reduction effects" against the Qwen version is the prerequisite for the next step the company names openly: training a future conductor on a domestically developed base model, aimed at customers who want sovereign control over the small orchestrator even as they keep pulling from foreign frontier models in the pool.
Two things the post does not put on the page. The evaluation set is described as custom-built and used only at test time, spanning knowledge, code correction, code generation, and graduate-level science, but no per-benchmark numbers or third-party comparisons are shown. And the reported cost saving is measured against a random-routing baseline, which is a floor rather than a competitive router. Both are reasonable for an internal validation write-up, but they mean the "equivalent" claim is a company self-report, not an independent bake-off.
If the method really is base-model-agnostic, the more interesting downstream move is the sovereign-conductor build the company flags for later: a small, controllable router placed in front of whichever frontier pool a buyer can legally reach.
Shared on Bluesky by 2 AI experts
Originally reported by sakana.ai
Read the original article →Original headline: Sakana AI