AI/TLDR

NVIDIA · 2026-08-24 · major

NVIDIA Groq 3 LPX — the agent inference chip enters full production

NVIDIA Groq 3 LPX is now in full production. The accelerator handles token generation for AI agents and reached 3,400 output tokens per second on Gemma 4 31B with a 100,000-token context. Nebius is the first cloud to adopt it.

NVIDIA Rubin rack family shown at Hot Chips 2026

NVIDIA's Groq 3 LPX accelerator is in full production, built to decode tokens fast enough to keep AI agents responsive.

Key specs

Output tokens/sec3,400
Responsiveness4x faster

Quick facts

MakerNVIDIA
StatusFull production
PlatformExtension of Vera Rubin NVL72
Job splitNVL72 reads context, LPX decodes tokens
Benchmark setupGemma 4 31B, 100K context (Artificial Analysis)
First cloudNebius, via Nebius Token Factory
Announced atHot Chips 2026

What is it?

Groq 3 LPX has entered full production — an NVIDIA accelerator built only for the token-generation half of AI inference. It extends the Vera Rubin NVL72 platform instead of replacing it, so a rack can run both parts of an agent's work. NVIDIA announced the milestone at Hot Chips 2026 on August 24, 2026.

How does it work?

Inference splits into two very different jobs, and Groq 3 LPX takes one of them. The Vera Rubin GPUs read and process long context, while the LPX accelerators decode the output tokens, which is the step where latency is felt. Running Gemma 4 31B with a 100,000-token context, Artificial Analysis measured 3,400 output tokens per second.

Why does it matter?

Agents spend most of their wall-clock time waiting on token decoding, so a chip aimed at that step changes how long a coding or research task takes. NVIDIA says Groq 3 LPX is 4x more responsive than the nearest alternative platform on latency-sensitive work. Nebius is the first AI cloud to adopt it, through its Nebius Token Factory inference service.

Who is it for?

AI infrastructure teams and inference providers

Frequently asked questions

When can I actually use Groq 3 LPX?
NVIDIA says Groq 3 LPX is now in full production, and Nebius is the first AI cloud to adopt it inside the Nebius Token Factory inference platform. NVIDIA's own release adds the usual caveat that products described are offered on a when-and-if-available basis, so there is no published general availability date for buying racks directly.
How is Groq 3 LPX different from Vera Rubin NVL72?
Groq 3 LPX and Vera Rubin NVL72 split one inference job in two. The NVL72 GPUs handle large-scale context processing — reading the prompt, documents and tool output. Groq 3 LPX handles token decoding, the latency-critical step that produces the answer. NVIDIA designed the LPX as an extension of the NVL72 platform, so the two run together.
Where does the Groq name in an NVIDIA product come from?
The Groq name comes from Groq Inc., the inference-chip startup whose hardware technology NVIDIA acquired in a $20 billion deal, as SiliconANGLE reports alongside the move of founder Jonathan Ross and president Sunny Madra to NVIDIA. Groq 3 LPX is the first NVIDIA-branded accelerator built on that LPU line.
What workloads is Groq 3 LPX meant for?
Groq 3 LPX targets latency-sensitive and agentic workloads — coding agents, chained tool calls and anything where a person or a downstream step is waiting on the next token. NVIDIA frames it as turning agentic tasks that took hours into ones that take minutes, and claims 4x better responsiveness than the nearest alternative platform.

Try it

Specs and platform details: https://www.nvidia.com/en-us/data-center/lpx/

Sources · 4 outlets

Tags

  • nvidia
  • groq-3-lpx
  • lpu
  • inference
  • ai-accelerator
  • hardware
  • vera-rubin
  • nvl72
  • agentic-ai
  • data-center
  • nebius
  • hot-chips-2026
  • ai-infrastructure

← All releases · Learn AI