sakana.ai web signal

Sakana AI ships Namazu, a Japanese-tuned Kimi K2.6 LLM

TL;DR

  • Sakana AI has turned Moonshot's open Kimi K2.6 into Namazu, a Japanese-specialized LLM tuned for business documents, keigo, and agent-style task work.
  • On Sakana's own comparison, Namazu lifts FairPoliticsQA from 34.10% to 56.30% while preserving Kimi's AIME26, MMLU-Pro, and LiveCodeBench v6 scores.
  • Pricing is $0.95 per million input tokens and $4.00 output on an OpenAI-compatible API, but Namazu is not available in the EU, UK, or Switzerland.

Sakana AI has turned Moonshot's open-weight Kimi K2.6 into Namazu, a Japanese-specialized LLM tuned for keigo, business documents, and the kind of long-running task work a Japanese office actually asks of a model. The pitch on Sakana AI's landing page is that the model preserves Kimi's underlying capabilities on AIME26, MMLU-Pro, and LiveCodeBench v6 while pulling markedly ahead on Japanese-specific benchmarks. The most eye-catching number is FairPoliticsQA, where Sakana reports a jump from 34.10% to 56.30% versus the Kimi K2.6 base.

The strategy is worth pausing on because it is the opposite of the frontier-lab move. Rather than train a giant model from scratch, Sakana takes strong open-weight foundations and post-trains them for Japanese language and culture. MarkTechPost reported that the Namazu series, first announced on March 24, 2026, has been applied to Kimi K2.6 alongside DeepSeek-V3.1-Terminus, Llama 3.1 405B, and gpt-oss-120B, and now powers Sakana Translate, the company's Japanese-English-Chinese translation and proofreading product.

Pricing looks positioned to compete rather than lead. Input runs $0.95 per million tokens, output $4.00, cached input $0.15, web search $7.00 per 1,000 calls, and code execution $0.12 per hour. The API is OpenAI-compatible, so for existing shops the switch is described as a base_url change rather than a rewrite. The model also does image recognition and can, per the landing page, search the web, write and run code, and carry complex tasks through to the end without a human in the loop.

The honest caveat is that the sharpest number here is a Sakana-run comparison against its own base model, and the landing page does not show head-to-head benchmarks against GPT, Gemini, or Claude on Japanese business tasks. Availability is a real constraint too: Namazu is currently unavailable in the EU, EEA, UK, and Switzerland while GDPR compliance work continues, so European buyers and Japanese groups with European subsidiaries are out of scope for now.

What is genuinely interesting is the shape of the bet. If a large chunk of Japanese enterprise value lives in workflows that require honorifics, business etiquette, and government or legal document handling, a locally-tuned open-weight model at these prices is a credible alternative to routing everything through a US frontier lab. The teams to watch are Japanese SIs and SaaS vendors who can point their existing OpenAI SDK code at Namazu and see whether the Japanese quality lift shows up in production.

Shared on Bluesky by 2 AI experts