AI/TLDR

Cerebras · 2026-08-18 · major

Cerebras CS-4 — three wafer-scale chips per rack, 30x GPU token speed

Cerebras CS-4 packs three Wafer Scale Engine 3 Turbo processors into one rack for 750 PFLOPS. Cerebras says it serves tokens up to 30 times faster per user than GPU systems and ships to first customers this quarter.

Cerebras CS-4 rack-scale AI system announcement

Cerebras CS-4 puts three wafer-sized AI chips in one rack and aims squarely at GPU inference speed.

Key specs

Ai compute per rack750 PFLOPS
Speed vs gpu systemsup to 30x

Quick facts

MakerCerebras
Processor3 x Wafer Scale Engine 3 Turbo (WSE-3T)
Transistors per wafer4 trillion
Cores per wafer900,000
On-wafer SRAM44 GB per wafer
Memory bandwidth129.6 PB/s per rack
AvailabilityFirst shipments this quarter

What is it?

The CS-4 is Cerebras' new rack-scale inference system, announced on August 18, 2026. It is built from three Wafer Scale Engine 3 Turbo processors, each a single wafer holding 4 trillion transistors, 900,000 AI cores and 44 GB of on-wafer SRAM across 46,225 square millimetres of silicon. One rack delivers 750 PFLOPS of AI compute.

How does it work?

Keeping weights in SRAM on the wafer removes the trip to external memory that limits GPU inference, and the CS-4 rack reaches 129.6 petabytes per second of memory bandwidth plus 7.2 terabits per second of system I/O. Direct Wafer Links join wafers with as little as two microseconds of latency, so clusters can hold models above 50 trillion parameters. The rack is the first machine on the Nexus architecture, which separates compute, power and I/O into swappable modules.

Why does it matter?

Speed per user, not total throughput, is what makes an agent feel usable, and agent loops burn tokens on reasoning and tool calls. Cerebras reports more than 4,400 tokens per second per user on GPT-OSS-120B and up to 10 times the throughput per watt of the CS-3, so the pitch is both faster responses and cheaper power. The Register frames it as another push at Nvidia's hold on inference hardware.

Who is it for?

inference providers and teams serving frontier-scale models

Frequently asked questions

When can you buy a Cerebras CS-4?
Cerebras says the first CS-4 shipments begin this quarter, which is the third quarter of 2026. The system is sold as a rack-scale solution rather than a single card, and Cerebras has not published a list price for it. The wafers are fabricated on TSMC's 5-nanometer process.
How is the CS-4 different from the CS-3?
The CS-4 is up to twice as fast as the CS-3 and delivers up to 10 times more throughput per watt, which is the number Cerebras leans on for data-center economics. It is also the first system built on the new Cerebras Nexus rack-scale architecture, which splits a rack into separate compute, power and I/O modules.
How large a model can a CS-4 cluster serve?
Cerebras states a CS-4 cluster supports models with more than 50 trillion parameters. Direct Wafer Links cut wafer-to-wafer latency to as low as two microseconds, which is what lets many racks act as one machine. On models above 10 trillion parameters, Cerebras reports more than 1,000 tokens per second.
How fast is the CS-4 on a real open model?
Cerebras reports more than 4,400 tokens per second per user on GPT-OSS-120B, which is the basis of its claim of up to 30 times the tokens-per-second-per-user of GPU systems. CTO Sean Lie argues that being 30 times faster "gives agentic systems room for significantly more reasoning and tool use".

Sources · 3 outlets

Tags

  • cerebras
  • cs-4
  • wse-3-turbo
  • nexus
  • wafer-scale
  • inference
  • hardware
  • accelerator
  • datacenter
  • tokens-per-second
  • agentic-inference

← All releases · Learn AI