Cerebras Runs OpenAI’s GPT-5.6 Sol at 750 Tokens Per Second in New Ultrafast Tier – Unite.AI

0
1
Cerebras Runs OpenAI’s GPT-5.6 Sol at 750 Tokens Per Second in New Ultrafast Tier – Unite.AI



Cerebras Runs OpenAI’s GPT-5.6 Sol at 750 Tokens Per Second in New Ultrafast Tier – Unite.AI

Cerebras is now running OpenAI’s flagship model at a speed no GPU cloud has publicly matched. On August 13, 2026, the wafer-scale chipmaker announced it powers GPT-5.6 Sol on a new OpenAI service tier called Ultrafast, delivering up to 750 output tokens per second and, by OpenAI’s account, running the model up to 14× faster than Standard processing. Ultrafast launches first in the OpenAI API as a limited preview for a select group of customers, with access expanding as capacity grows.

The claim at the center is a specific one: frontier intelligence without the speed penalty. GPT-5.6 Sol is OpenAI’s most capable model, and on Cerebras silicon it generates tokens fast enough to sit inside real-time products rather than behind an overnight batch job. Both companies frame the tier as removing the tradeoff between a model smart enough for high-stakes work and one fast enough to use while the work is still happening.

The Speed Numbers Behind Ultrafast

The 750 output tokens per second figure is the headline, but the more telling comparisons are the head-to-head runs Cerebras published. Against output speeds reported by Artificial Analysis, the company says GPT-5.6 Sol on Ultrafast runs 11× faster than Claude Fable 5 and 5× faster than Opus 4.8 on Fast mode.

Cerebras also put the tier through a full pass of Humanity’s Last Exam, a 2,500-question benchmark pitched at PhD-level difficulty. GPT-5.6 Sol on Ultrafast answered all 2,500 questions in 11 hours and 11 minutes; Claude Fable 5 needed 78 hours and 27 minutes to reach comparable conclusions: nearly 7× slower, in Cerebras’s telling. On GDP-Val, a benchmark for economically valuable knowledge work, the company reports a 5.6× end-to-end speedup over Standard processing with no quality degradation.

These are vendor-run evaluations, and Cerebras is explicit about that: the Humanity’s Last Exam comparison was benchmarked by Cerebras on July 10 and July 13–15, 2026, and the GDP-Val figure comes from its own July 31, 2026 testing. Treat them as the company’s own measurements, not independent results.

Why the Wafer-Scale Chip Wins on Latency

The mechanism matters here, because it explains why a relatively small chipmaker is serving OpenAI’s biggest model at speeds the GPU incumbents haven’t matched. Fast inference on a large model is fundamentally a data-movement problem: on GPUs, model weights must be shuttled repeatedly between on-chip memory and off-chip storage to generate each successive token, and memory bandwidth becomes the bottleneck.

Cerebras’s answer is to eliminate that movement. Its Wafer-Scale Engine packs 44 GB of SRAM onto a single wafer-sized chip, so the model’s weights stay on-chip and tokens flow through layers pipelined across wafers without interruption. Because the weights never leave the silicon, the approach scales with model size — which is the company’s argument that the speed advantage holds as frontier models grow. For inference economics, that is the whole game: the cost and latency of serving a model are dominated by how fast you can feed weights to the compute, and keeping 44 GB resident on one die attacks that directly.

A $10 Billion Partnership Reaches the Flagship

Ultrafast is the most visible product yet of a relationship that has been building for months. OpenAI tapped Cerebras for $10 billion in low-latency compute earlier in 2026, and this launch puts that capacity behind the company’s top model rather than a smaller or specialized one. OpenAI describes Ultrafast as “the next step” in the partnership to bring ultra-low-latency inference to its platform.

For Cerebras, the placement is significant. The startup has long argued its wafer-scale architecture is the right shape for inference even as the market’s center of gravity sits with GPU suppliers — a contest playing out across the accelerator business as incumbents move to bake models directly into their silicon. Landing the serving layer for OpenAI’s flagship gives Cerebras a production reference account at the top of the market. The model itself anchors the GPT-5.6 family OpenAI launched (Sol as the flagship, alongside the balanced Terra and the cost-efficient Luna), so Ultrafast attaches Cerebras to the front of that lineup.

Who Gets It and What It’s For

OpenAI is positioning the tier at time-sensitive, high-stakes work: incident response while an outage is still unfolding, financial research while market conditions are moving, real-time customer support and voice, commerce, and live research loops that used to run overnight. Early access has gone to companies across coding, commerce, and finance, including Jane Street, Podium, Basis, and Rogo.

> “The increase in speed brought by Cerebras is impressive,” said John Crepezzi, AI Assistants at Jane Street, in OpenAI’s announcement. “It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them.”

OpenAI is keeping the rollout narrow on purpose. The company says it is using the preview period to learn where an order-of-magnitude speed change creates the most value, and will expand access as capacity grows. Both OpenAI and Cerebras are taking sign-ups for updates as the preview widens.

What Happens Next

The near-term observable is capacity, not capability. Ultrafast is a limited preview, and both companies tie any broader availability to capacity growth rather than a fixed date — so the pace of expansion is the thing to watch. The longer question is whether Cerebras’s on-chip-memory advantage holds as OpenAI’s models scale, which is exactly the bet the wafer-scale architecture is built on.