Seraphic and CrowdStrike: A New Chapter for Browser Security

Read more
close icon

September 29, 2026

World Action Models and the Shift Towards Predictive Robotics

Type

Deep Dives

Contributors

Murat Kilicoglu

Over the past few years, robotics has made meaningful progress on a problem that once looked unsolvable: training a single model to interpret a visual scene, understand a natural-language instruction, and turn that understanding into physical action. Vision-language-action models, or VLAs, are the clearest expression of this shift. Systems such as RT-2, OpenVLA, and Physical Intelligence’s π0 build on large vision-language models and extend them into the action space, allowing a robot to generalize across a wider range of objects, instructions, and environments than earlier task-specific systems. This is an important development, but it still leaves open a more basic question about physical intelligence. Knowing what action is appropriate in a given situation is not the same as understanding what that action is likely to cause.

An interesting example is a simple household task such as placing a mug in a sink. A capable VLA may correctly identify the mug, understand the instruction, and generate a sequence of motor commands that usually works. The difficulty is that physical environments are rarely as clean as the demonstrations used to train the policy. The mug may be partially blocked, the handle may face the wall, another object may be in the way, or the first grasp may slip. At that point, the robot needs more than semantic knowledge and a learned mapping from images to actions. It needs some representation of how the scene is likely to change as a result of what it does. This is the basic motivation behind the recent interest in world action models.

From world models to world action models

The term “world action model,” or WAM, is relatively new and the category is still fairly loose. A 2026 paper uses the term to describe models that combine predictive state modeling with action generation, although several closely related systems use different terminology. For now, it is probably more useful to focus on the underlying idea than on where exactly the category begins and ends.

A VLA is primarily trained to map an observation and an instruction to an action. A world model tries to represent how the environment evolves, often conditioned on a possible action. A world action model brings those two together, so that a prediction of the future informs the action the robot ultimately takes.

This distinction also helps clarify why a strong video model is not automatically a useful world model for robotics. A generative model can produce a visually reasonable sequence showing an object being picked up or moved without having a reliable model of which robot command would produce that outcome in the current scene. The generated motion may look right while getting the geometry, timing, or contact dynamics slightly wrong. In robotics, those differences are critical. A plausible video of a glass moving across a table is not the same as understanding how a particular gripper, at a particular pose, should interact with that glass without knocking it over.

Exhibit: Conceptual comparison of the input-output formulations of vision-language-action models, world action models, and standard world models, highlighting WAM’s capability to jointly predict actions and future observations.

The broader idea has a much longer history. Richard Sutton’s Dyna work in 1990 combined direct experience with planning using a learned model of the environment. Later research extended this idea to high-dimensional environments. World models showed that an agent could learn a compressed representation of visual dynamics and train a controller entirely inside that learned environment. PlaNet learned latent dynamics directly from images and planned within that latent state. Dreamer used imagined trajectories inside a learned world model to train its policy and value function, while MuZero showed that useful planning does not necessarily require reconstructing every detail of an observation. A model can be valuable even if it represents only the aspects of the environment that are relevant to the decision.

That last idea is particularly important for robotics. A robot does not need to predict the exact texture of a countertop or the reflection on the side of a stainless-steel sink. It may, however, need a good estimate of whether a mug will collide with another object, whether a grasp is likely to remain stable, or whether a partially completed action has moved the scene into a state that requires a different plan. The central design question is therefore not simply whether a robot should predict the future, but which representation of the future is sufficient to improve control.

What changed to make world action models possible

Recent robot foundation models emerged from a different research lineage. RT-1, RT-2, OpenVLA, π0, and related systems were designed primarily as generalist policies. They combine visual observations, language, and large-scale robot data to produce actions directly. RT-2 was particularly notable because it demonstrated that concepts learned from internet-scale vision-language pre-training could transfer into robotic behavior, while OpenVLA showed that a large open model could be trained on nearly a million robot demonstrations. These systems significantly improved the semantic side of robotic control by allowing policies to reuse representations learned far beyond the narrow set of objects and tasks present in an individual robot dataset.

What they do not automatically provide is an explicit model of physical consequences. A model may recognize an object as a mug and understand that mugs can be placed in sinks, yet still have only a weak representation of how a particular grasp will affect that mug. The distinction is subtle because a sufficiently capable policy may implicitly learn some dynamics as part of predicting good actions. There is no sharp boundary between a reactive policy with strong internal representations and a system with a dedicated world model. The current interest in WAMs is partly an attempt to make this predictive part of physical intelligence. A capable physical intelligence system also needs some way of representing what its actions are likely to do to the world around it.

times
#
# #