Cerebras · 2026-08-18 · major
Cerebras CS-4 — three wafer-scale chips per rack, 30x GPU token speed
Cerebras CS-4 packs three Wafer Scale Engine 3 Turbo processors into one rack for 750 PFLOPS. Cerebras says it serves tokens up to 30 times faster per user than GPU systems and ships to first customers this quarter.

Cerebras CS-4 puts three wafer-sized AI chips in one rack and aims squarely at GPU inference speed.
Key specs
| Ai compute per rack | 750 PFLOPS |
|---|---|
| Speed vs gpu systems | up to 30x |
Quick facts
| Maker | Cerebras |
|---|---|
| Processor | 3 x Wafer Scale Engine 3 Turbo (WSE-3T) |
| Transistors per wafer | 4 trillion |
| Cores per wafer | 900,000 |
| On-wafer SRAM | 44 GB per wafer |
| Memory bandwidth | 129.6 PB/s per rack |
| Availability | First shipments this quarter |
What is it?
The CS-4 is Cerebras' new rack-scale inference system, announced on August 18, 2026. It is built from three Wafer Scale Engine 3 Turbo processors, each a single wafer holding 4 trillion transistors, 900,000 AI cores and 44 GB of on-wafer SRAM across 46,225 square millimetres of silicon. One rack delivers 750 PFLOPS of AI compute.
How does it work?
Keeping weights in SRAM on the wafer removes the trip to external memory that limits GPU inference, and the CS-4 rack reaches 129.6 petabytes per second of memory bandwidth plus 7.2 terabits per second of system I/O. Direct Wafer Links join wafers with as little as two microseconds of latency, so clusters can hold models above 50 trillion parameters. The rack is the first machine on the Nexus architecture, which separates compute, power and I/O into swappable modules.
Why does it matter?
Speed per user, not total throughput, is what makes an agent feel usable, and agent loops burn tokens on reasoning and tool calls. Cerebras reports more than 4,400 tokens per second per user on GPT-OSS-120B and up to 10 times the throughput per watt of the CS-3, so the pitch is both faster responses and cheaper power. The Register frames it as another push at Nvidia's hold on inference hardware.
Who is it for?
inference providers and teams serving frontier-scale models
Frequently asked questions
- When can you buy a Cerebras CS-4?
- Cerebras says the first CS-4 shipments begin this quarter, which is the third quarter of 2026. The system is sold as a rack-scale solution rather than a single card, and Cerebras has not published a list price for it. The wafers are fabricated on TSMC's 5-nanometer process.
- How is the CS-4 different from the CS-3?
- The CS-4 is up to twice as fast as the CS-3 and delivers up to 10 times more throughput per watt, which is the number Cerebras leans on for data-center economics. It is also the first system built on the new Cerebras Nexus rack-scale architecture, which splits a rack into separate compute, power and I/O modules.
- How large a model can a CS-4 cluster serve?
- Cerebras states a CS-4 cluster supports models with more than 50 trillion parameters. Direct Wafer Links cut wafer-to-wafer latency to as low as two microseconds, which is what lets many racks act as one machine. On models above 10 trillion parameters, Cerebras reports more than 1,000 tokens per second.
- How fast is the CS-4 on a real open model?
- Cerebras reports more than 4,400 tokens per second per user on GPT-OSS-120B, which is the basis of its claim of up to 30 times the tokens-per-second-per-user of GPU systems. CTO Sean Lie argues that being 30 times faster "gives agentic systems room for significantly more reasoning and tool use".