Google TPUv7 Ironwood beats Nvidia B200 on inference cost
TL;DR
- Google's TPUv7 Ironwood hit $0.181 per million tokens versus $0.222 on Nvidia's B200 and $0.276 on B300 in SemiAnalysis's FP8 test.
- SemiAnalysis says Anthropic is now the biggest TPU user, with commitments over one million units surpassing DeepMind by 2029.
- Google's TorchTPU native PyTorch backend is expected to leave private beta and open source around October.
Google's TPUv7 Ironwood delivered $0.181 per million tokens in SemiAnalysis's InferenceX benchmark, against $0.222 on Nvidia's B200 and $0.276 on the B300, at 100 tokens per second per user on FP8 models. SemiAnalysis frames the chip as offering "up to 50% better performance per dollar" versus those two Nvidia parts.
At a lower 20 tokens-per-second concurrency, Ironwood put up 9,364 tokens per chip per second to B200's 8,903, a 5% edge in raw throughput on top of the price gap.
The other half of the story is software. Google is moving off its previous TorchAX translation layer to what SemiAnalysis calls a "native PyTorch backend," a stack called TorchTPU that lets developers address TPUs as first-class PyTorch devices. "The stack is expected to leave private beta & be open sourced around October," the report says, naming Inferact, RadixArk, and Red Hat as outside collaborators alongside vLLM and SGLang integrations. The bring-up model for testing was Qwen3.5 397B in FP8.
Each Ironwood chip is now "two separate compute dies, each running its own independent logical device," with "2 TensorCores and 4 third-generation SparseCores." A successor part, TPUv8i, will use a new "Boardfly" topology that SemiAnalysis says cuts network diameter "by more than 50% versus a similarly sized 3D torus."
Buried in the customer section is the loudest claim: "Anthropic being the biggest user of TPUs, surpassing Deepmind's own use by 2029," with commitments the report puts at more than one million units across direct purchases and Google Cloud rental. The InferenceX numbers are SemiAnalysis's own testing, not independent third-party benchmarks.
Shared on Bluesky by 1 AI expert
Originally reported by newsletter.semianalysis.com
Read the original article →Original headline: TPU Inference Externalization Full Steam Ahead - InferenceX