Cerebras powers OpenAI's Ultrafast mode for GPT-5.6 Sol at 750 tokens per second
Cerebras Systems is the hardware behind OpenAI's newly announced Ultrafast mode for GPT-5.6 Sol, running the model at speeds the two companies say reach 750 output tokens per second with no quality tradeoff. Cerebras detailed the partnership in a blog post published August 13, 2026, the same day OpenAI's API changelog announced Ultrafast mode as a limited preview.
What's new
Cerebras states the core claim directly: "Cerebras powers GPT-5.6 Sol on Ultrafast mode, delivering up to 750 output tokens per second and without any quality compromise." The company frames the speedup relative to other frontier models rather than only against OpenAI's own Standard tier, claiming "GPT-5.6 Sol on Ultrafast mode runs 11x faster than Fable 5, and 5x faster than Opus 4.8 on Fast mode" — both Anthropic models.
On benchmark throughput, Cerebras cites Humanity's Last Exam, a 2,500-question evaluation: "GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes" versus "Claude Fable 5 needed 78 hours and 27 minutes" for a comparable run — roughly a 7x wall-clock difference at comparable accuracy.
Context
The speed comes from Cerebras's wafer-scale architecture rather than conventional GPU clusters. As the company describes it: "we pack 44 GB of SRAM on each wafer-sized chip. Weights stay on-chip, and tokens flow uninterrupted through model layers pipelined across wafers." That design keeps model weights resident in on-chip memory rather than shuttling them across a network between GPUs, and pipelines inference across multiple wafer-scale chips built for exactly this workload rather than general-purpose GPU-to-GPU networking.
This is Cerebras's most prominent frontier-lab inference deal yet, putting its wafer-scale inference architecture behind the largest model in OpenAI's current lineup rather than smaller open-weight models, which has been its more typical positioning. Ultrafast mode itself remains a limited preview for select OpenAI customers, with no committed general-availability date or public pricing from OpenAI.
Why it matters
For Cerebras, powering a flagship OpenAI model's fastest inference tier is a significant validation of wafer-scale inference at the frontier-model scale, not just for smaller or open-weight models where the approach has previously been demonstrated. It's a concrete data point in the broader argument that specialized inference silicon can meaningfully outperform GPU clusters on latency-sensitive workloads.
For OpenAI, partnering with a non-NVIDIA inference specialist for its highest-speed tier — while its data-center buildout elsewhere leans heavily on NVIDIA GPUs — signals that different parts of the stack are being optimized for different hardware, rather than standardizing on one vendor end-to-end. If Ultrafast mode moves from preview to general availability at scale, it would be one of the highest-profile production deployments of Cerebras hardware to date.
Corroborating sources
- Cerebras
https://www.cerebras.ai/blog/accelerating-gpt-5-6-sol-ultrafast-with-openai
“Cerebras powers GPT-5.6 Sol on Ultrafast mode, delivering up to 750 output tokens per second and without any quality compromise”