newsletter.semianalysis.com web signal

OpenAI Details Jalapeño, Its First Broadcom-Built Inference ASIC That Beats Nvidia Rubin on Perf-per-Watt

Summary

SemiAnalysis published a deep dive on OpenAI's first custom inference chip Jalapeño, taped out with Broadcom in just 16 months on TSMC N3P. The B0 stepping hits 13.4 PFLOPs of MXFP4 at 700W (vs Rubin's 900-1,150W), pairs HBM4 at 15.4TB/s, and posts 700+ tok/s/user on DeepSeek R1 and ~1,400 tok/s/user on GPT-OSS. The Verge separately reports OpenAI benchmarks put Jalapeño at 1.5-1.9x more work per watt than Nvidia across GPT-OSS, DeepSeek R1 and Kimi K2.5 1T.

Shared on Bluesky by 4 AI experts