World Action Models for Robot Manipulation
Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blog

Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models | NVIDIA Technical Blog

6/15/2026

What this post added

This post analyzes the recent rise of World-Action Models (WAMs) in robot foundation models, contrasting them with Vision-Language-Action (VLA) models. It details the core hypotheses for WAMs, such as addressing the language-to-action grounding gap by leveraging pretrained video or world-model backbones that already model scene dynamics. The post categorizes modern WAMs based on what the model predicts (inverse dynamics, joint prediction, representation-only) and how actions are integrated (default action tokens, action as image, latent actions/plans). It also discusses architectural considerations and hypothesizes why WAMs have gained prominence recently, suggesting a potential paradigm shift towards WAMs or hybrid VLA/WAM approaches for generalist robot policies.

Read the original post ↗