OpenAI Details Jalapeño, Its First Broadcom-Built Inference ASIC That Beats Nvidia Rubin on Perf-per-Watt
Summary
SemiAnalysis published a deep dive on OpenAI's first custom inference chip Jalapeño, taped out with Broadcom in just 16 months on TSMC N3P. The B0 stepping hits 13.4 PFLOPs of MXFP4 at 700W (vs Rubin's 900-1,150W), pairs HBM4 at 15.4TB/s, and posts 700+ tok/s/user on DeepSeek R1 and ~1,400 tok/s/user on GPT-OSS. The Verge separately reports OpenAI benchmarks put Jalapeño at 1.5-1.9x more work per watt than Nvidia across GPT-OSS, DeepSeek R1 and Kimi K2.5 1T.
Shared on Bluesky by 4 AI experts
Originally reported by newsletter.semianalysis.com
Read the original article →Original headline: OpenAI Details Jalapeño, Its First Broadcom-Built Inference ASIC That Beats Nvidia Rubin on Perf-per-Watt