Apple · 2026-08-25 · major
Apple M6 and M5 Ultra — 2nm silicon and 512GB for on-device LLMs
Apple's M6 is its first 2nm chip, and the new M5 Ultra bonds four dies over UltraFusion to reach an 80-core GPU with 512GB of unified memory. Apple says a Mac Studio can now run frontier-class LLMs entirely on device.

Apple's first 2nm chip and its first quad-die M-series part, both aimed squarely at running big models on your desk.
Quick facts
| Maker | Apple |
|---|---|
| M6 | First 2nm chip; 12-core CPU, 12-core GPU, dual 16-core Neural Engine |
| M6 memory | Up to 32GB unified at 170GB/s |
| M5 Ultra | Quad-die UltraFusion; up to 36-core CPU, 80-core GPU |
| M5 Ultra memory | Up to 512GB unified at 1.2TB/s |
| Mac mini (M6) price | From $899 |
| Availability | Pre-order August 25, 2026; ships September 22 |
Pricing
| Mac Studio (M5 Max) · $2,299 education | $2,499 |
|---|---|
| Mac Studio (M5 Ultra) · $5,099 education | $5,499 |
What is it?
Apple M6 is the company's first chip built on a 2-nanometer process, pairing a 12-core CPU and 12-core GPU with a dual 16-core Neural Engine. Alongside it, the M5 Ultra is the first M-series part to join four dies into one chip, scaling to a 36-core CPU, an 80-core GPU and 512GB of unified memory. The two land in a new Mac mini and a new Mac Studio.
How does it work?
The quad-die trick is UltraFusion, Apple's packaging link, which connects two dual-die M5 Max chips at over 4.4TB/s between dies with six times the connection density of the previous generation. Both chips put Neural Accelerators inside each GPU core and feed them from one pool of unified memory — 170GB/s on M6, 1.2TB/s on M5 Ultra — so weights do not have to be copied between CPU, GPU and Neural Engine.
Why does it matter?
A 512GB unified memory ceiling is the headline for local inference: Apple says Mac Studio lets users run massive models entirely on device with complete privacy, without counting tokens or watching cloud bills. Apple also reports that four Mac Studio machines clustered together deliver up to 3x faster AI inference than one. For M6, nearly 30 percent more peak GPU compute for AI over M5 mostly shows up as quicker prompt processing on on-device LLMs.
Who is it for?
people running local models, ML engineers, Mac developers
Frequently asked questions
- Can a Mac Studio run large language models locally?
- Apple states that Mac Studio with M5 Ultra lets users run massive models entirely on device with complete privacy, without counting tokens or worrying about rising cloud costs. The 512GB unified memory ceiling is what makes frontier-class model sizes fit. Apple also reports that a cluster of four Mac Studio systems delivers up to 3x faster AI inference than a single machine.
- How much faster is M6 than M5 for AI work?
- Apple claims M6 delivers nearly 30 percent more peak GPU compute for AI than M5, up to 1.2x faster multithreaded CPU performance, and a dual 16-core Neural Engine with up to 2x the peak compute of previous generations. Apple frames the GPU gain as improving prompt processing speed when running on-device LLMs rather than as a general throughput number.
- What is UltraFusion and why does M5 Ultra use four dies?
- UltraFusion is Apple's die-to-die packaging link, and M5 Ultra is the first M-series chip to use it in a quad-die arrangement, joining two dual-die M5 Max chips into one part. The connection carries over 4.4TB/s between dies with six times the connection density of the prior generation, which is how M5 Ultra reaches 36 CPU cores and 80 GPU cores.
- When do the new Macs actually ship?
- Pre-orders for the new Mac mini with M6 and the new Mac Studio with M5 Max or M5 Ultra opened on August 25, 2026, and both machines ship to customers from September 22, 2026. Apple notes one exception: the 512GB unified memory configuration of Mac Studio arrives later, in late October 2026.
Try it
Pre-orders are open now at apple.com; both Macs ship September 22, 2026, with the 512GB configuration following in late October.