hardmaru

Why they matter

Founder with public evidence across AI business.

AI signals
12
past 30d
Sources
7
distinct domains
Discussions
0
past 30d
Latest signal
1d ago
View every signal from hardmaru →
Co-Founder & CEO, Sakana AI 🎏 → @sakanaai.bsky.social Visit → https://sakana.ai/

Articles & links

Read the full open-access paper: www.nature.com/articles/s41... Blog: sakana.ai/smart-cellul... Congratulations to the team on this achievement!

Smart cellular bricks for decentralized shape classification and damage recovery | Nature Communications nature.com
AI Weekly's analysis →
  • Cubic bricks running identical neural cellular automata policies classified four 3D shapes with 98.97% accuracy in simulation and 100% on physical hardware.
  • Physical builds ranged from 26 bricks for a guitar to 197 for a round table, converging on a shape label in fewer than 60 update cycles.
  • The same decentralized framework detects structural damage with over 90% accuracy and guides regrowth by predicting one of six axis directions.
Read full analysis →
View on Bluesky · ♥ 10 ↻ 2 ↩ 0 · 3 from the directory shared this · 77d ago
↻ hardmaru reposted
Sakana AI @sakanaai.bsky.social

Introducing "Scaling In-Context Imitation Learning" (SAIL) to be presented at #IROS2026. This work is a collaboration between Sakana AI and the University of Tokyo. Blog: pub.sakana.ai/sail Paper: arxiv.org/abs/2603.08269 What does a robot need before it can tackle a new task?…

SAIL: Test-Time Scaling for In-Context Imitation Learning with VLM arxiv.org
AI Weekly's analysis →
  • SAIL reframes robot imitation as MCTS over full trajectories, guided by a VLM scorer, step-level feedback, and a retrieval archive of past successes.
  • Across six simulation manipulation tasks, average success rose from 25% at one rollout to 73% at 45 MCTS nodes, reaching 95% on HandOverBanana.
  • On a real-world BlockIntoBowl task with a LeRobot SO-101 arm, the method succeeded in five of six trials; the paper is accepted to IROS 2026.
Read full analysis →
View on Bluesky →

I am incredibly proud of our Tokyo team for shipping this. By orchestrating the world’s models, we are delivering the resilient blueprint required for AI sovereignty. Read our full vision and results here: sakana.ai/fugu-release 🐡

Sakana AI sakana.ai
View on Bluesky · ♥ 12 ↻ 0 ↩ 0 · 6 from the directory shared this · 98d ago

“World Models in Natural and Artificial Intelligence” brings together pioneers including Douglas Hofstadter, Michael Levin, Josh Tenenbaum, Samuel Gershman, and Melanie Mitchell to ask: What if the next leap in AI requires not just more data, but systems that model themselves?

royalsocietypublishing.org
View on Bluesky · ♥ 4 ↻ 0 ↩ 1 · 5 from the directory shared this · 16d ago
↻ hardmaru reposted
@tksiia.bsky.social

Excited to share CoffeeBench!!☕️☕️☕️ We evaluate LLM agents in a 90-day B2B coffee supply-chain economy spanning farmers, roasters, and retailers, where autonomous firms negotiate, manage inventory, set prices, handle invoices, and manage cash flow. arxiv.org/abs/2606.16613 gi…

CoffeeBench: Benchmarking Long-Horizon LLM Agents in Heterogeneous Multi-Agent Economies arxiv.org
AI Weekly's analysis →
  • CoffeeBench runs LLMs as autonomous business operators across a 90-day, six-firm economic simulation.
  • Claude Haiku 4.5 exhibits 'idle-drift': producing coherent plans but persistently choosing inaction.
  • Higher-performing models communicate more actively with counterpart firms, correlating with better outcomes.
Read full analysis →
View on Bluesky →

By orchestrating the world's models, we are building the resilient infrastructure required for AI sovereignty. Try: sakana.ai/fugu Blog: sakana.ai/fugu-max-rel... 🐡

Sakana Fugu — Multi-agent System as A Model sakana.ai
AI Weekly's analysis →
  • Fugu routes tasks through a dynamic multi-agent pipeline exposed as a single OpenAI-compatible API, removing orchestration setup from users.
  • The system draws on two ICLR 2026 papers: TRINITY assigns Thinker/Worker/Verifier roles; Conductor uses reinforcement learning to design coordination strategies.
  • Fugu Ultra scored 73.7 on SWE Bench Pro and 93.2 on LiveCodeBench; base Fugu reached 95.5 on GPQA-D, per Sakana's own benchmark reporting.
Read full analysis →
View on Bluesky · ♥ 5 ↻ 0 ↩ 0 · 3 from the directory shared this · 17d ago

By orchestrating the world's models, we are building the resilient infrastructure required for AI sovereignty. Try: sakana.ai/fugu Blog: sakana.ai/fugu-max-rel... 🐡

Sakana AI sakana.ai
AI Weekly's analysis →
  • Fugu Max prices at $2 per million input tokens and $6 per million output tokens, which Sakana claims runs 40-60% below Sonnet 5, GPT 5.6 Terra, and Kimi K3.
  • The system routes each task across a swappable pool of open-weight and specialized models, with NVIDIA's Nemotron folded in via an August 2026 collaboration.
  • Fugu Ultra v2 scores 48.3 on Chartography against Opus 5's 27.3, and does so without Fable 5, Fable 5.1, or GPT-6-Astra in its agent pool.
Read full analysis →
View on Bluesky · ♥ 5 ↻ 0 ↩ 0 · 3 from the directory shared this · 17d ago
↻ hardmaru reposted
Sakana AI @sakanaai.bsky.social

Introducing PC-ALM: a local-learning alternative to backprop that trains 1000-layer neural nets using only local dynamics. Blog: pub.sakana.ai/pc-alm/

Augmented Lagrangian Predictive Coding: training 1000-layer networks without backpropagation pub.sakana.ai
AI Weekly's analysis →
  • PC-ALM adds a per-layer Lagrange multiplier to predictive coding so that only layer-local updates recover backpropagation-aligned credit signals.
  • Sakana reports 1000-layer residual MLPs on MNIST land within roughly two percentage points of backprop using a T=2L inference budget.
  • On Fashion-MNIST at width 32 and depth 32, PC-ALM hits 77.75% accuracy versus 78.66% for backprop and 68.13% for vanilla predictive coding.
Read full analysis →
View on Bluesky →

Read the full open-access paper: www.nature.com/articles/s41... Blog: sakana.ai/smart-cellul... Congratulations to the team on this achievement!

Smart Cellular Bricks: Towards Collective Intelligence for the Physical World sakana.ai
AI Weekly's analysis →
  • IT University of Copenhagen, Sakana AI, and Autodesk built cubic bricks that classify their own assembled 3D shape using only neighbor-to-neighbor communication.
  • In simulation the system hit 98.97% accuracy across 500+ bricks; four physical objects (26 to 197 bricks) all reached correct consensus in under 60 cycles.
  • The same Neural Cellular Automata substrate detects local damage at 94.8% average accuracy, with some shapes degrading only minimally at 15% brick failure.
Read full analysis →
View on Bluesky · ♥ 10 ↻ 2 ↩ 0 · 3 from the directory shared this · 77d ago

Dive into our new AI Picbreeder Experiment here: pub.sakana.ai/picbreeder-v...

The AI Picbreeder Experiment: In Search of Automatic Open-Endedness pub.sakana.ai
AI Weekly's analysis →
  • Researchers from NYU, MIT and Sakana AI replaced Picbreeder's human selectors with frontier VLMs and reported clear qualitative gaps versus the historical human baseline.
  • Gemini-2.5-pro topped the models tested, but archives showed 'mode collapse' without exploratory noise and 'auto-sycophantic' loops when given more history.
  • Scaling to 1,000 agents widened coverage yet 10 to 20 percent of the archive turned into uninterpretable adversarial psychedelic patterns.
Read full analysis →
View on Bluesky · ♥ 6 ↻ 1 ↩ 0 · 3 from the directory shared this · 80d ago

Recent commentary

Language models and coding agents are great, but there is more to life, and more to AI, than just LLM agents.

View on Bluesky · ♥ 50 ↻ 5 ↩ 4 · 77d ago

In hardmaru's orbit

Center = hardmaru. Left = members they follow (green edges). Right = members who follow them (blue edges). Top = mutual follows (orange edges, slightly larger). Drag any node to reposition; click to open that profile.

Are you hardmaru? Show it.

Add the Who’s Who of AI badge to your site or bio. It links back to this profile.

Listed in AI Weekly's Who's Who of AI

Markdown: [![Listed in AI Weekly's Who's Who of AI](https://aiweekly.co/modules/custom/aiweekly_whoswho/images/whoswho-badge.svg)](https://aiweekly.co/whos-who/person/hardmaru-bsky-social)