OpenAI Previews GPT-5.6 Sol Ultrafast at 14x Speed on Cerebras
TL;DR
- GPT-5.6 Sol on Ultrafast mode runs up to 14 times faster than Standard processing and generates up to 750 output tokens per second.
- The tier is in limited preview through the OpenAI API, powered by Cerebras under an ultra-low-latency inference partnership between the two companies.
- OpenAI is using Ultrafast internally for incident response, reading logs, analyzing traces, and helping prepare or validate fixes.
OpenAI has quietly put its fastest inference tier yet into a limited preview, and the numbers, if they hold in production, rework the latency budget for anything interactive built on the model. According to Help Net Security, GPT-5.6 Sol on Ultrafast mode runs up to 14 times faster than Standard processing and generates up to 750 output tokens per second, delivered through the OpenAI API to a select group of customers and powered by Cerebras under the two companies' partnership on ultra-low-latency inference.
Preview customers are reportedly testing Ultrafast across coding, commerce, financial research, support, and other interactive applications in production environments. John Crepezzi, described in the piece as AI Assistants at Jane Street, said the speed increase from Cerebras 'enables different ways of using the models' and lets developers work 'in a more focused and productive way alongside them.' OpenAI is also eating its own cooking. Internally the mode is being used for incident response, spanning reading logs, analyzing traces, synthesizing conversations, identifying follow-up checks, and helping prepare or validate fixes. Research teams that would normally launch experiments overnight and review results the next morning can instead complete multiple iterations during the workday, the company said.
That shift is what matters for anyone designing on top of frontier models. When a response returns in a second instead of ten, product patterns change: agent loops can plan and self-check more times per user turn, coding assistants stop feeling like batch jobs, and interfaces built around a 'typing' pause become interfaces built around instant answers. It is also the second time in three months our tracker has logged a Cerebras throughput claim tied to a specific frontier model, after May's trillion-parameter benchmark run, which suggests specialty silicon is now part of how OpenAI segments its own product line rather than a curiosity on the side.
A few gaps to sit with. Help Net Security does not disclose pricing, capacity, or a date for general availability, and the 14x and 750 tokens per second figures are OpenAI's own numbers rather than independent measurements. There is also no word on context length limits or how throughput holds up under real concurrency, which is where wafer-scale systems have historically been touchy. Treat the ceiling as a demo ceiling until customers outside the preview publish their own numbers.
If the mode graduates on similar performance, the biggest winners are teams building agentic and interactive products where wall-clock latency was the actual ship blocker, and Cerebras itself, which now has an OpenAI reference for its architecture against the Nvidia default most inference budgets still assume.
Originally reported by helpnetsecurity.com
Read the original article →Original headline: OpenAI Previews Ultrafast Mode for GPT-5.6 Sol, 14x Faster on Cerebras at 750 Tokens/Sec