█

AI/TLDR

NVIDIA · 2026-10-07 · major

Long-WAM — NVIDIA's robot world-action model remembers 19 seconds of video

Long-WAM is a world-action model from NVIDIA, MIT, HKU and UCSD that gives robots a long visual memory and still runs in real time. It scores 99.5% on LIBERO-Long, and its code and checkpoints are open under Apache 2.0.

Long-WAM project preview from NVIDIA Research

A robot policy that watches the last 19 seconds, imagines what comes next, then acts, fast enough for real robots.

Key specs

Libero long99.5%
Latency (rtx 5090)107.4 ms / action chunk

Quick facts

MakersNVIDIA, MIT, HKU, UCSD
PaperarXiv 2610.10528 (Oct 7, 2026)
History lengthUp to 19.2 s of visual context
BenchmarksLIBERO-Long 99.5%, RoboTwin 2.0 94.4%, DOMINO 34.9%
Robots testedUnitree G1, YAM arm
Runs onRTX 5090, DGX Spark, Jetson AGX Thor
LicenseApache 2.0 (code + checkpoints)

What is it?

Long-WAM is a framework for scaling how much visual history a causal world-action model can use when controlling a robot, while keeping inference real-time. NVIDIA, MIT, HKU and UCSD posted the paper on October 7, 2026, with code in the LongLive repository and checkpoints on Hugging Face. Its motto is "Remember the past. Imagine the future. Act."

How does it work?

The method first pretrains an autoregressive video model on robot and egocentric video with no action labels, so it learns to predict the future from the past. It then adapts that model to output both future video latents and chunks of robot actions. Streaming observation encoding, asynchronous execution and hardware-specific acceleration bring one action chunk down to 107.4 ms on an RTX 5090, including the future-video prediction.

Why does it matter?

The authors find that longer history helps far more when the base video model was pretrained autoregressively than bidirectionally, where it shows no net gain. That matters for long or dynamic tasks: the policy succeeded 19 of 20 times at dynamic cup stacking, where two baseline models failed every trial. Used as the executor under a GPT-6 Astra planner, success on 50 RoboCasa365 tasks rose from 31.4% to 54.4%.

Who is it for?

robotics and embodied-AI researchers

Frequently asked questions

Are the Long-WAM weights open?
Yes. NVIDIA publishes Long-WAM checkpoints under the Efficient-Large-Model organisation on Hugging Face, including policies for LIBERO, RoboTwin 2.0, RoboCasa365 and the YAM arm, five RoboCasa GR-1 history variants and two LongLive 2.0 Robot video generators. The code and checkpoints are released under Apache 2.0; third-party parts keep their own licenses.
What hardware can run Long-WAM?
The Long-WAM authors deploy the model on an RTX 5090, a DGX Spark and a Jetson AGX Thor. On the RTX 5090 one action chunk takes 107.4 ms including future-video latent prediction. The README targets Python 3.10 or newer with CUDA PyTorch 2.7.x for the deployment environment and uses a separate environment for training.
Can I run Long-WAM on a real robot today?
Long-WAM ships real-robot interfaces for the YAM arm, Franka and Unitree G1, but the README says deployment defaults to read-only inference that sends no robot commands. Deployment checkpoint paths are blank, so users must supply their own policy and a calibrated driver. The README also marks full benchmark reproduction as not yet verified.
How is Long-WAM related to LongLive?
Long-WAM lives in the Long-WAM folder of NVIDIA's LongLive repository, which already hosts the LongLive long-video generation stack. Long-WAM builds on that work: it pretrains a LongLive 2.0 Robot video model on robot and egocentric footage, then adapts it to predict robot actions alongside future video.

Try it

python scripts/download_checkpoint.py robocasa365 --output checkpoints/robocasa365

Sources · 4 outlets

Tags

  • nvidia
  • long-wam
  • longlive
  • world-action-model
  • world-model
  • robotics
  • embodied-ai
  • video-model
  • unitree-g1
  • open-weights
  • apache-2-0

← All releases · Learn AI