
Beyond VLAs: How World Action Models Reshape Robot Manipulation | NVIDIA Technical Blog
8/4/2026
This post introduces the concept of World Action Models (WAMs) as a successor to Vision-Language Action (VLA) models for robot manipulation. It highlights the limitations of VLAs in physical generalization due to their focus on semantic description rather than world dynamics. WAMs, by building on video world models, learn physical dynamics, leading to improved generalization, reduced data requirements for adaptation, and better performance on new robots. The post details how the NVIDIA Cosmos 3 model, an omni-model world foundation model, provides a robust foundation for WAMs. It also presents Cosmos3-Nano-Policy-DROID and Cosmos3-Edge-Policy-DROID as examples of post-trained WAMs for robot policies, showcasing their ability to imagine while acting and their deployment considerations for workstation serving and on-device inference.


