AI/TLDR

NVIDIA · 2026-06-01 · major

NVIDIA Nemotron 3 Ultra — 550B Mamba-Transformer Mixture-of-Experts With 55B Active Parameters Tops US Open-Weights Leaderboard at 48 on Artificial Analysis Intelligence Index, Ships June 4 on Hugging Face, OpenRouter, ModelScope, and build.nvidia.com

Jensen Huang's Computex keynote unveiled NVIDIA's flagship open-weights model: 550B total / 55B active Mamba-Transformer MoE serving 300+ tokens/sec with a 1M-token context window. Lands June 4 across HF, OpenRouter, and NIM microservices.

Artificial Analysis Intelligence Index chart with Nemotron 3 Ultra leading US open-weights models
Artificial Analysis

NVIDIA's biggest open-weights model: a 550B Mamba-Transformer MoE that beats every other US open model on intelligence and inference speed.

Key specs

Active params55B
Context window1M tokens
Total parameters550B
Sparsity~90%
Artificial analysis intelligence index48
Inference speed (deep infra)300+ tok/s
Inference speedup vs comparable open models5x
Cost reduction vs comparable open models~30%

What is it?

Nemotron 3 Ultra is the flagship of NVIDIA's Nemotron 3 family, post-trained for long-running agents across coding, research, and enterprise workflows. It targets the same workloads as Claude Opus and GPT-5.5 but ships with open weights, NVFP4 quantization, and a 1M-token context window.

How does it work?

Ultra uses a hybrid Mamba-Transformer mixture-of-experts: Mamba blocks handle long sequences with linear-cost attention while Transformer MoE layers route tokens to ~10% of the 550B parameters per step. It's natively post-trained for Hermes Agent, LangChain Deep Agents, OpenClaw, OpenHands, and OpenCode harnesses.

Why does it matter?

At 48 on the Artificial Analysis Intelligence Index, Ultra is now the smartest US open-weights model — well ahead of Gemma 4 31B (39) and gpt-oss-120b (33), but still trailing China's Kimi K2.6 (54). At 300+ tokens/sec on DeepInfra it also beats DeepSeek V4 Pro and Kimi K2.6 (50–100 tok/s) for latency-sensitive agent loops.

Who is it for?

agent builders and inference platforms that want open weights with frontier intelligence

Try it

Available June 4 on Hugging Face, ModelScope, OpenRouter, and build.nvidia.com NIM microservices

Sources · 5 outlets

Tags

  • nvidia
  • nemotron
  • open-weights
  • mixture-of-experts
  • mamba-transformer
  • agentic
  • computex-2026
  • nvfp4
  • long-context

← All releases · Learn AI