hardmaru

Why they matter

Founder with public evidence across AI business.

AI signals
6
past 30d
Sources
2
distinct domains
Discussions
2
past 30d
Latest signal
2d ago
View every signal from hardmaru →
Co-Founder & CEO, Sakana AI 🎏 → @sakanaai.bsky.social Visit → https://sakana.ai/

Articles & links

Read the full open-access paper: www.nature.com/articles/s41... Blog: sakana.ai/smart-cellul... Congratulations to the team on this achievement!

Smart cellular bricks for decentralized shape classification and damage recovery | Nature Communications nature.com
AI Weekly's analysis
  • Cubic bricks running identical neural cellular automata policies classified four 3D shapes with 98.97% accuracy in simulation and 100% on physical hardware.
  • Physical builds ranged from 26 bricks for a guitar to 197 for a round table, converging on a shape label in fewer than 60 update cycles.
  • The same decentralized framework detects structural damage with over 90% accuracy and guides regrowth by predicting one of six axis directions.
Read full analysis →
View on Bluesky · ♥ 10 ↻ 2 ↩ 0 · 3 from the directory shared this · 36d ago

I am incredibly proud of our Tokyo team for shipping this. By orchestrating the world’s models, we are delivering the resilient blueprint required for AI sovereignty. Read our full vision and results here: sakana.ai/fugu-release 🐡

Sakana AI sakana.ai
View on Bluesky · ♥ 12 ↻ 0 ↩ 0 · 6 from the directory shared this · 57d ago
hardmaru reposted
@tksiia.bsky.social

Excited to share CoffeeBench!!☕️☕️☕️ We evaluate LLM agents in a 90-day B2B coffee supply-chain economy spanning farmers, roasters, and retailers, where autonomous firms negotiate, manage inventory, set prices, handle invoices, and manage cash flow. arxiv.org/abs/2606.16613 gi…

CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies arxiv.org
AI Weekly's analysis
  • CoffeeBench runs LLMs as autonomous business operators across a 90-day, six-firm economic simulation.
  • Claude Haiku 4.5 exhibits 'idle-drift': producing coherent plans but persistently choosing inaction.
  • Higher-performing models communicate more actively with counterpart firms, correlating with better outcomes.
Read full analysis →
View on Bluesky →

Read the full open-access paper: www.nature.com/articles/s41... Blog: sakana.ai/smart-cellul... Congratulations to the team on this achievement!

Smart Cellular Bricks: Towards Collective Intelligence for the Physical World sakana.ai
AI Weekly's analysis
  • IT University of Copenhagen, Sakana AI, and Autodesk built cubic bricks that classify their own assembled 3D shape using only neighbor-to-neighbor communication.
  • In simulation the system hit 98.97% accuracy across 500+ bricks; four physical objects (26 to 197 bricks) all reached correct consensus in under 60 cycles.
  • The same Neural Cellular Automata substrate detects local damage at 94.8% average accuracy, with some shapes degrading only minimally at 15% brick failure.
Read full analysis →
View on Bluesky · ♥ 10 ↻ 2 ↩ 0 · 3 from the directory shared this · 36d ago

Today, we are taking a major step toward that future with the launch of Sakana Fugu. Fugu dynamically orchestrates the world’s best models to tackle complex tasks. We’re proving that a well-orchestrated pool of swappable agents can match restricted frontier models like Fable/M…

Sakana Fugu — Multi-agent System as A Model sakana.ai
AI Weekly's analysis
  • Fugu routes tasks through a dynamic multi-agent pipeline exposed as a single OpenAI-compatible API, removing orchestration setup from users.
  • The system draws on two ICLR 2026 papers: TRINITY assigns Thinker/Worker/Verifier roles; Conductor uses reinforcement learning to design coordination strategies.
  • Fugu Ultra scored 73.7 on SWE Bench Pro and 93.2 on LiveCodeBench; base Fugu reached 95.5 on GPQA-D, per Sakana's own benchmark reporting.
Read full analysis →
View on Bluesky · ♥ 4 ↻ 1 ↩ 1 · 2 from the directory shared this · 57d ago

I’ve been thinking about how agents can learn inside world models for years. We decided to scale up our RSI Lab to bridge recursive self-improvement with physical AI and robotics. We are looking for frontier researchers and engineers to join us in Tokyo, Japan: sakana.ai/caree…

Sakana AI sakana.ai
AI Weekly's analysis
  • Sakana AI is hiring a Member of Technical Staff for its Tokyo-based Recursive Self-Improvement (RSI) Lab, targeting researchers frustrated with brute-force scaling.
  • The role spans four tracks including world models as verifiable simulators for agentic reasoning and open-ended evolutionary dynamics applied to algorithmic domains.
  • Sakana points to prior work as receipts: Darwin Gödel Machine on SWE-bench, ALE-Agent winning AtCoder Heuristic Contest 058, ShinkaEvolve with 150 samples.
Read full analysis →
View on Bluesky · ♥ 19 ↻ 3 ↩ 1 · 2 from the directory shared this · 7d ago

Sakana Namazu: An LLM API with Japanese-vibes! 🎏 Built for Japanese enterprises, featuring frontier-level reasoning and built-in agentic tools. Blog post: sakana.ai/namazu-api#E...

Sakana AI sakana.ai
AI Weekly's analysis
  • Sakana AI released the Sakana Namazu API, an OpenAI-compatible endpoint for the LLM that has been powering its Sakana Chat product.
  • The model is built on Moonshot AI's open Kimi K2.6, tuned in-house on proprietary data for Japanese language and business context.
  • Sakana reports FairPoliticsQA rising from 34.10% to 56.30% versus the base model, with additional gains claimed on JFBench, AIME26, and coding.
Read full analysis →
View on Bluesky · ♥ 12 ↻ 2 ↩ 0 · 2 from the directory shared this · 15d ago
hardmaru reposted
Sakana AI @sakanaai.bsky.social

🐟 Sakana Namazu API 公開 🐟 本日、Sakana AIは大規模言語モデル「Namazu」をアップデートし、API「Sakana Namazu(サカナ・ナマズ)」として提供を開始しました。 Sakana Namazu API: sakana.ai/namazu 🐟

Sakana Namazu — 日本の感性で、考え、動くLLM sakana.ai
AI Weekly's analysis
  • Sakana AI has turned Moonshot's open Kimi K2.6 into Namazu, a Japanese-specialized LLM tuned for business documents, keigo, and agent-style task work.
  • On Sakana's own comparison, Namazu lifts FairPoliticsQA from 34.10% to 56.30% while preserving Kimi's AIME26, MMLU-Pro, and LiveCodeBench v6 scores.
  • Pricing is $0.95 per million input tokens and $4.00 output on an OpenAI-compatible API, but Namazu is not available in the EU, UK, or Switzerland.
Read full analysis →
View on Bluesky →

Dive into our new AI Picbreeder Experiment here: pub.sakana.ai/picbreeder-v...

The AI Picbreeder Experiment: In Search of Automatic Open-Endedness pub.sakana.ai
AI Weekly's analysis
  • Researchers from NYU, MIT and Sakana AI replaced Picbreeder's human selectors with frontier VLMs and reported clear qualitative gaps versus the historical human baseline.
  • Gemini-2.5-pro topped the models tested, but archives showed 'mode collapse' without exploratory noise and 'auto-sycophantic' loops when given more history.
  • Scaling to 1,000 agents widened coverage yet 10 to 20 percent of the archive turned into uninterpretable adversarial psychedelic patterns.
Read full analysis →
View on Bluesky · ♥ 6 ↻ 1 ↩ 0 · 2 from the directory shared this · 39d ago

Recent commentary

Language models and coding agents are great, but there is more to life, and more to AI, than just LLM agents.

View on Bluesky · ♥ 50 ↻ 5 ↩ 4 · 36d ago

In hardmaru's orbit

Center = hardmaru. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.