OpenAI previews Ultrafast mode, running GPT-5.6 Sol up to 14x faster on Cerebras
OpenAI opened a limited preview of Ultrafast mode on August 13, 2026, a new API service tier that runs GPT-5.6 Sol at up to 14 times the speed of standard processing by routing inference through Cerebras' wafer-scale hardware.
What's new
OpenAI describes the new tier directly: "Ultrafast, a new service tier that runs GPT‑5.6 Sol up to 14× faster than Standard processing." In practice that means "Ultrafast generates up to 750 output tokens per second" for GPT-5.6 Sol requests routed to the new tier — several times the throughput most production frontier-model deployments reach today.
Access is limited for now: "GPT‑5.6 Sol on Ultrafast mode is available in a limited preview today to a select group of customers. We'll expand access as capacity grows." OpenAI has not published pricing for the tier alongside the preview announcement. The company frames the release as a continuation of existing work with its hardware partner: "Ultrafast marks the next step in our partnership with Cerebras to bring ultra-low-latency inference to OpenAI's platform."
Context
OpenAI and Cerebras have been building toward this for months. GPT-5.6 Sol itself launched in July running on Cerebras' wafer-scale WSE-3 chips, which eliminate the interconnect bottleneck that limits how fast a cluster of separate GPUs can stream tokens by putting memory and compute on a single silicon wafer. That earlier deployment was already reported to reach roughly 750 tokens per second in some configurations — figures that now line up with what OpenAI is describing as the ceiling for the new Ultrafast tier specifically, suggesting Ultrafast is the productized, generally-offered version of capability that was previously more experimental or restricted.
The move continues a broader industry pattern: OpenAI, like Google and Amazon before it, is increasingly pairing its own frontier models with specialized inference silicon — Cerebras, Groq, and SambaNova have all built businesses around exactly this kind of ultra-low-latency serving — rather than relying solely on general-purpose GPU clusters for every workload.
Why it matters
Inference speed has become a genuine product differentiator as more usage shifts toward latency-sensitive, real-time applications: voice agents, live coding assistance, and multi-step agentic workflows all degrade noticeably when token generation lags behind what a user or downstream system needs. A 14x speedup, even in limited preview, gives OpenAI a concrete answer to competitors who have marketed speed as their primary edge — Cerebras and Groq in particular have built go-to-market strategies almost entirely around raw tokens-per-second numbers.
It also signals that OpenAI's Cerebras partnership is deepening rather than staying a one-off deployment for a single model. If Ultrafast expands beyond the initial "select group of customers" and OpenAI publishes pricing that makes it viable for production use, it would give developers a genuine choice between cost-optimized and latency-optimized serving within OpenAI's own platform — rather than needing to route latency-sensitive workloads to a separate inference provider running open-weight models instead.
Corroborating sources
- Openai
https://openai.com/index/previewing-ultrafast
“Ultrafast, a new service tier that runs GPT‑5.6 Sol up to 14× faster than Standard processing”