newsletter.semianalysis.com web signal

Cerebras CS-4 posts 4,400 tokens/sec, up to 30x vs GPUs

TL;DR

  • Cerebras CS-4 delivers more than 4,400 tokens per second per user on GPT-OSS-120B, up to 30 times a GPU inference baseline.
  • Each rack pairs three WSE-3 Turbo wafers for 750 PFLOPS, 129.6 PB/s of memory bandwidth, and two-microsecond wafer-to-wafer latency.
  • AMD Instinct and AWS Trainium handle prompt processing while CS-4 acts as the decode accelerator; first shipments begin this quarter.

Cerebras's new CS-4 posts more than 4,400 tokens per second per user on OpenAI's GPT-OSS-120B, a mark the company puts at up to 30 times the throughput of the fastest GPU-based inference service today, per SemiAnalysis. The system was unveiled August 18, 2026, with first shipments this quarter.

"In AI, speed is productivity. Historically, fast inference meant using smaller models," said Andrew Feldman, CEO and co-founder of Cerebras. CTO Sean Lie framed the pitch around agents: "Being 30 times faster gives an agentic system room for significantly more reasoning."

Each rack packs three WSE-3 Turbo wafers for 750 PFLOPS of AI compute, 129.6 petabytes per second of aggregate memory bandwidth, and 7.2 Tbps of system I/O. Wafer-to-wafer latency is quoted as low as two microseconds, and Cerebras says the platform can support models with more than 50 trillion parameters. The Register puts power draw at roughly 120 to 140 kW per rack.

The system is not a full-stack alternative to a GPU cluster. Cerebras offloads prompt processing to AWS Trainium and AMD Instinct chips, positioning CS-4 mainly as the decode accelerator that turns the resulting state into tokens. Dylan Patel of SemiAnalysis, quoted in the launch, said the machine "will scale ultrafast tokens for larger models and significant user volumes."

Shared on Bluesky by 1 AI expert