github.com via Hacker News

openTPU Ships AI-Developed FPGA Accelerator for Qwen3, LFM2.5

TL;DR

  • openTPU is a full-stack open-source inference accelerator with RTL, ISA, compiler, simulator and host tools, targeting a Xilinx Kintex-7 FPGA on an Inspur PCIe card.
  • The author, FeSens, labels the project 'developed by AI,' framing it as a test of whether agents can design the chip that runs their own inference.
  • Reported throughput on LFM2.5-230M is 59.0 tok/s decode and 295.6 tok/s prefill in int8, with DRAM running 82-94% of a 17.1 GB/s peak.

FeSens's openTPU repository ships a full-stack inference accelerator covering SystemVerilog RTL, a custom instruction set, a bit-exact Python simulator, a kernel compiler, a profiler called Lens, and host utilities including otpu-chat. It targets a Xilinx Kintex-7 xc7k480t FPGA on an Inspur YPCB-00338 PCIe card with 4 GiB of DDR3-1066 memory across two channels. The README's tagline is blunt: "An open-source AI accelerator, developed by AI."

The design is set up to run small language models on the board, including LFM2.5-230M, several Qwen3 and Qwen3.5 variants up to 4B, SmolLM3-3B, Phi-4-mini, Gemma 4 E2B and E4B, and two mixture-of-experts configurations (LFM2.5-8B-A1B and Qwen3.5-35B-A3B). On LFM2.5-230M in int8, the author reports 59.0 tok/s decode and 295.6 tok/s prefill; the 4-bit path reaches 85.8 tok/s decode and 335.4 tok/s prefill. DRAM utilization during decoding runs 82-94% of the board's 17.1 GB/s peak.

The framing is a methodology experiment more than a product. "openTPU brings the lessons of auto-arch-tournament to AI accelerators," the README says. "It asks two questions: how far can AI agents go at hardware design, and can they build the chip that runs their own inference?" The design keeps nothing opaque to make the answer legible: "There is no cache and no hidden scheduling: every data movement is an instruction, so a trace shows exactly where the cycles go."

The README does not quantify how much of the RTL, ISA, or compiler came from agents versus human hands, nor name the agent harness or models that produced the design. The pitch is pedagogical: "If you want to understand how an AI accelerator works, from a matmul in Python down to the wires, this is a good place to start."