
6/15/2026
What this post added
This post analyzes the recent rise of World-Action Models (WAMs) in robot foundation models, contrasting them with Vision-Language-Action (VLA) models. It details the core hypotheses for WAMs, such as addressing the language-to-action grounding gap by leveraging pretrained video or world-model backbones that already model scene dynamics. The post categorizes modern WAMs based on what the model predicts (inverse dynamics, joint prediction, representation-only) and how actions are integrated (default action tokens, action as image, latent actions/plans). It also discusses architectural considerations and hypothesizes why WAMs have gained prominence recently, suggesting a potential paradigm shift towards WAMs or hybrid VLA/WAM approaches for generalist robot policies.