AI/TLDR

OpenAI · 2026-08-13 · major

Ultrafast mode — GPT-5.6 Sol at 750 tokens per second on Cerebras

Ultrafast mode is a new OpenAI API service tier that runs GPT-5.6 Sol at up to 750 output tokens per second, up to 14 times faster than standard. Cerebras wafer-scale chips power it. Limited preview for now.

Cerebras and OpenAI artwork for GPT-5.6 Sol Ultrafast mode

OpenAI's new API tier runs GPT-5.6 Sol on Cerebras hardware at up to 750 output tokens per second.

Key specs

Output speed750 tokens/sec
Gdp val speedup5.6x end-to-end

Quick facts

MakerOpenAI, hardware by Cerebras
ModelGPT-5.6 Sol
Peak output speed750 tokens/sec
Speed vs standardUp to 14x faster
QualitySame as GPT-5.6 Sol Standard
HardwareCerebras Wafer-Scale Engine, 44 GB on-chip SRAM
AvailabilityLimited preview on the OpenAI API, waitlist open

What is it?

Ultrafast mode is a new service tier in the OpenAI API that runs GPT-5.6 Sol at up to 750 output tokens per second, up to 14 times faster than the standard tier. OpenAI says the model keeps the same intelligence as GPT-5.6 Sol Standard, so only the speed changes. It opened on August 13, 2026 as a limited preview for selected customers.

How does it work?

Cerebras Wafer-Scale Engine hardware supplies the speed. Inference on a large model is bottlenecked by moving weights between memory layers, so the Cerebras chip keeps 44 GB of SRAM on-chip and the weights stay reachable without repeated transfers. Cerebras measured the result at 11 times faster than Fable 5 and 5 times faster than Opus 4.8 running in Fast mode.

Why does it matter?

Latency, not intelligence, is what blocks many production uses of a frontier model. On GDP-Val, a benchmark of economically valuable knowledge work, Ultrafast mode delivered a 5.6x end-to-end speedup with no drop in quality. OpenAI points at voice, customer support, commerce, developer agents, financial research, and security response, and says it already uses the tier internally to read logs and traces during incidents.

Who is it for?

Teams building latency-sensitive agents, voice, and support products

Frequently asked questions

How much does Ultrafast mode cost?
OpenAI has not published a price for Ultrafast mode. Both the OpenAI preview post and the Cerebras announcement of August 13, 2026 describe the speed, the hardware, and the preview access process without listing a per-token rate. Businesses joining the waitlist submit workload details, latency requirements, and expected usage instead of picking a published tier.
Who can access Ultrafast mode right now?
Ultrafast mode is a limited preview restricted to selected OpenAI API customers, with access widening over time. Businesses that want in join a waitlist and submit their workload details, latency requirements, and expected usage. There is no self-serve toggle in the API for general accounts during the preview period.
Is Ultrafast mode less accurate than standard GPT-5.6 Sol?
No. OpenAI and Cerebras both state that GPT-5.6 Sol on Ultrafast mode runs with the same intelligence as GPT-5.6 Sol Standard. On GDP-Val the tier produced a 5.6x end-to-end speedup with no quality degradation, so the trade is purely latency for hardware availability rather than accuracy.
Which workloads does OpenAI recommend Ultrafast mode for?
OpenAI points Ultrafast mode at live and near-production work where latency matters as much as model intelligence: voice, customer support, commerce, developer agents, financial research, and security response. OpenAI also says it uses the tier internally to analyze logs and traces during incidents, and to compress overnight research cycles into several iterations inside one workday.

Sources · 3 outlets

Tags

  • openai
  • cerebras
  • gpt-5-6-sol
  • ultrafast
  • inference
  • api
  • latency
  • wafer-scale
  • agents
  • tool
  • serving

← All releases · Learn AI