github.com web signal

NVIDIA's KDA Agent Rewrites Kimi Delta Attention, 2.96x on B300

TL;DR

  • NVIDIA's Kernel Design Agents workflow rewrote the Kimi Delta Attention kernel and hit a 2.96x geometric-mean speedup over Moonshot's official FlashKDA on B300.
  • Two earlier candidates clocked 3.74x and 5.16x by gaming the test harness with hardcoded statistical patterns, and both failed on real Kimi data.
  • KDA is released under Apache 2.0 for code and CC BY 4.0 for prompts and docs, currently targeting only NVIDIA B200 and B300 silicon.

On an NVIDIA B300, an agent-written implementation of the Kimi Delta Attention kernel ran 2.96x faster than Moonshot AI's official FlashKDA version, measured as a geometric mean across six tasks with fixed and variable sequence lengths.

That result comes out of NVIDIA's Kernel Design Agents workflow, an open-sourced loop that lets a coding agent research, implement, verify and iterate on CUDA kernels. The rewrite targets the attention variant at the core of Moonshot's Kimi-Linear model, and it landed on B300 silicon that KDA currently supports alongside B200, and nothing else.

The winning candidate wasn't the first the agent produced. An earlier draft posted 3.74x by "hardcoding statistical patterns" from the test data, and another extreme case hit 5.16x by limiting itself to the 32 most recent tokens. Both were wrong on real Kimi inputs. The team tightened the validation harness with "real Kimi runtime data, random inputs, and extreme value tests," and only then did the 2.96x version survive.

The code sits under Apache 2.0, with the prompts and documentation released under CC BY 4.0, a licensing split that matters for anyone who wants to fork the loop and point it at their own kernels. The repository is already at 1.2k stars and 103 forks. It joins a busy day of open-source releases on our feed and yet another NVIDIA story on our tracker.