This website uses cookies

Read our Privacy policy and Terms of use for more information.

TL;DR: VLAs turn visual and language understanding into robot actions. World Action Models (WAMs) also learn how the world changes, using predicted futures to guide actions. Cascaded WAMs separate prediction from control, while joint WAMs learn both together. Early results are promising, but WAMs are not consistently better than VLAs and often require more inference compute.

The idea is that learning how the world changes can help robots choose actions in unfamiliar situations.

World models are bringing AI closer to understanding how the physical world works. But what is that understanding for? Can we stop at predicting what happens next? The next step is to put that knowledge into practice – and that means taking action.

A relatively new class of models is trying to connect these two abilities. They are called World Action Models (WAMs), and they aim to give robots not only a way to predict how the world changes, but also a way to act based on those predictions.

But we already have Vision-Language-Action models (VLAs), one of the main approaches in Physical AI, combining visual perception, language understanding, and robot control. So why do we need something else?

That’s exactly what we’ll explore today. We’ll look at how WAMs work, what makes them different from VLAs, the main types of WAM architectures, and whether predicting the future actually helps robots perform better.

Another question is whether robots should learn to act directly from what they see or learn to imagine what their actions will do first.

Let's explore this new path from understanding the world to interacting with it – it is very interesting! Like peaking into the future a little bit.

In today’s episode:

  • What is a World Action Model—and where did the term come from?

  • VLA vs. WAM: how their approaches differ

  • Cascaded WAMs: predicting a future, then turning it into actions

  • Joint WAMs: learning future states and actions together

  • How DreamZero connects video prediction with robot control

  • How video models represent robot commands

  • Do WAMs generalize better than VLAs?

  • Should robots react or imagine?

Not interested in this topic? Watch our overview of the freshly released Open d1 – a decision model family from Liquid AI (and why it might be better than Jev) →

and now to the main topic:

What is a World Action Model?

A World Action Model is a model that looks deeper into physics and connects a model of how the world changes with a model of what actions a robot or an agent should take.

To be a WAM, a model needs two things:

  • A prediction of the future. This might be an image, video, optical flow, a point cloud, or a learned latent representation – in other words, a compact numerical description of the predicted state.

  • Action generation tied to that prediction. The model can predict the future first and use it to derive actions, or generate future states and actions jointly.

When the term WAM appeared

While NVIDIA’s 2026 paper “World Action Models are Zero-shot Policies” popularized the term to distinguish these architectures from VLAs, the concept evolved across several earlier works. Researchers occasionally used the phrase throughout 2024 and 2025 to describe models unifying world dynamics with action generation, with the earliest notable mentions appearing by October 2024.

If you ask yourself: Was it Jurgen Schmidhuber who invented the term, we can tell you that: Schmidhuber helped lay the foundations for today’s world-action models, with work on neural world models and controllers dating back to 1990.

But as far as we know, the story goes this way:

On October 1, 2024 Jie Cheng and colleagues published a paper called “Scaling Offline Model-Based RL via Jointly-Optimized World-Action Model Pretraining”, where they introduced JOWA model. They use World-Action Model to describe a model that learns how an environment changes and which actions to take. Trained on Atari games, JOWA uses a shared transformer backbone jointly optimized with world-modeling and temporal-difference (TD) losses to learn environment dynamics and action values.

Then, on March 21, 2025 “DyWA: Dynamics-adaptive World Action Model for Generalizable Non-prehensile Manipulation” paper was published by Jiangran Lyu and colleagues at Peking University and Galbot. It uses World Action Model in a robotics setting as a policy model that jointly predicts robot actions and their corresponding future states. Future-state prediction is used as an additional training signal to improve action learning.

And only in 2026 the most popular paper about WAM came out. NVIDIA research team presented DreamZero on February 17, 2026 in their “World Action Models are Zero-shot Policies”. DreamZero is a WAM built on a pretrained video diffusion model that jointly generates future video frames and actions. To be precise, it learns to predict what a robot’s actions will look like and which motor commands will produce them – together.

Model Architecture of DreamZero

Image Credit: World Action Models are Zero-shot Policies research paper

The project leads included Seonghyeon Ye, Yuke Zhu, Jim Fan, and Joel Jang. So it makes sense to associate WAMs with Jim Fan and NVIDIA in the context of DreamZero, but it is not accurate to say they were first to coin the term.

NVIDIA is positioning DreamZero WAM as an alternative to VLAs, especially when it comes to learning new movements and physical skills. Since we’ve touched on this debate, let’s unpack what is actually different about VLAs and WAMs.

VLA vs. WAM: The Difference in Ideas and Principles

While the boundary between them is blurring, VLAs and WAMs differ fundamentally in architecture and foundations. That distinction starts with what each model learns:

  • A VLA learns: “What action or action sequence should I take?”

  • A WAM learns: “What will happen in the world, and which action will lead to the future I want?”

VLA

(Vision-Language-Action Models)

WAM

(World Action Models)

Main question

What should I do?

What will happen if I do this? + What should I do?

Typical foundation

Vision-Language model (VLM)

World/video model

Primary learning focus

Semantics → action

Dynamics + action

Output

Robot action or action sequence

Action + future state/representation, or an action informed by learned world dynamics

Strength

Generalization across language and meaning

Physical and spatiotemporal dynamics

Analogy

Perceive → act

Imagine → act

So let’s break down the workflow distinctions on the simplest example which is practical for a robot. The task is: “Pick up the mug and put it in the sink.” →

You need to read it to really understand the difference.

Learn the basics and go deeper with us. Find inspiration for what to build and the knowledge to put it into practice.

Join Premium members from top companies like Microsoft, NVIDIA, Google, HF, OpenAI, a16z, plus AI labs such as Ai2, MIT, Berkeley, .gov, and thousands of others to really understand what’s going on in AI. 

FAQ

What Is a World Action Model (WAM)?

A World Action Model is an AI model that connects predictions about how the world changes with action generation. It can predict future states and use them to generate robot commands, or learn to predict future states and actions jointly. The key idea is that predictions about the world help the model decide what to do.

What Is the Difference Between a VLA and a World Action Model?

A Vision-Language-Action model typically uses visual observations and language instructions to generate robot actions. A World Action Model additionally learns how the environment changes and connects those predictions to action generation. In simple terms, a VLA focuses on what action to take, while a WAM also considers what will happen as a result.

What Are the Main Types of World Action Models?

There are two broad types: cascaded WAMs and joint WAMs. Cascaded WAMs use a world model to predict future states and a separate controller to turn those predictions into actions. Joint WAMs learn future-state prediction and action generation together. Both types can use different representations, including video, visual tokens, and compact latent states.

Are World Action Models Better Than VLAs?

Not consistently. Recent evaluations show that WAMs can outperform VLAs on some robotics benchmarks, while VLAs perform better on others. Results depend on the task, training data, and environmental changes. WAMs can also require more inference time because they predict future states alongside actions. There is not yet enough evidence to call WAMs universally superior.

Do World Action Models Need to Generate Video?

No. Some WAMs, such as DreamZero, generate future video and robot actions together. Others predict compact latent representations, motion features, or structured states instead of full video frames. What makes a model a WAM is not video generation itself, but the connection between predicting future world states and generating actions.

Reply

Avatar

or to participate

Keep Reading

View more
caret-right