NVIDIA · 2026-06-01 · major
NVIDIA Nemotron 3 Ultra — 550B Mamba-Transformer Mixture-of-Experts With 55B Active Parameters Tops US Open-Weights Leaderboard at 48 on Artificial Analysis Intelligence Index, Ships June 4 on Hugging Face, OpenRouter, ModelScope, and build.nvidia.com
Jensen Huang's Computex keynote unveiled NVIDIA's flagship open-weights model: 550B total / 55B active Mamba-Transformer MoE serving 300+ tokens/sec with a 1M-token context window. Lands June 4 across HF, OpenRouter, and NIM microservices.

NVIDIA's biggest open-weights model: a 550B Mamba-Transformer MoE that beats every other US open model on intelligence and inference speed.
Key specs
| Active params | 55B |
|---|---|
| Context window | 1M tokens |
| Total parameters | 550B |
| Sparsity | ~90% |
| Artificial analysis intelligence index | 48 |
| Inference speed (deep infra) | 300+ tok/s |
| Inference speedup vs comparable open models | 5x |
| Cost reduction vs comparable open models | ~30% |
What is it?
Nemotron 3 Ultra is the flagship of NVIDIA's Nemotron 3 family, post-trained for long-running agents across coding, research, and enterprise workflows. It targets the same workloads as Claude Opus and GPT-5.5 but ships with open weights, NVFP4 quantization, and a 1M-token context window.
How does it work?
Ultra uses a hybrid Mamba-Transformer mixture-of-experts: Mamba blocks handle long sequences with linear-cost attention while Transformer MoE layers route tokens to ~10% of the 550B parameters per step. It's natively post-trained for Hermes Agent, LangChain Deep Agents, OpenClaw, OpenHands, and OpenCode harnesses.
Why does it matter?
At 48 on the Artificial Analysis Intelligence Index, Ultra is now the smartest US open-weights model — well ahead of Gemma 4 31B (39) and gpt-oss-120b (33), but still trailing China's Kimi K2.6 (54). At 300+ tokens/sec on DeepInfra it also beats DeepSeek V4 Pro and Kimi K2.6 (50–100 tok/s) for latency-sensitive agent loops.
Who is it for?
agent builders and inference platforms that want open weights with frontier intelligence
Try it
Available June 4 on Hugging Face, ModelScope, OpenRouter, and build.nvidia.com NIM microservices