World model

LLM foundationsModels and inferencePublished By Simon Budziak

A world model is an AI system that learns an internal simulation of an environment and uses it to predict what happens next. Instead of predicting the next word the way a language model does, it models objects, space and cause and effect, which lets an agent plan actions before taking them.

Why are AI labs building world models?

Because next-token prediction runs out of world. An LLM has read about gravity but never watched anything fall, which caps how well it can plan physical or spatial work. A world model gives an agent a place to rehearse: it can simulate an action internally, check the predicted outcome and only then act. That matters for robotics, autonomous operations and any computer use task where a wrong click has consequences, and it is why several frontier labs now describe world models as the next capability jump after reasoning models.

How does a world model actually work?

The core loop is encode, predict, compare. The model compresses each observation, a video frame or a sensor reading, into an internal state, then predicts how that state changes when an action is taken. Training pushes the prediction toward what actually happened, so over millions of transitions the model absorbs the dynamics of its environment: what moves, what collides, what an intervention changes. Prediction conditioned on actions is the defining ingredient, because it turns the model into something an agent can query. Try action A in imagination, read the predicted outcome, compare it with action B, then commit. The imagined attempt costs almost nothing; the real one, in physical AI especially, does not.

How is a world model trained and evaluated?

Two families dominate. Generative approaches predict the next observation directly, pixels included, which makes results easy to inspect but spends most of the model’s capacity on visual detail that rarely matters for a decision. Latent approaches predict in the compressed state space instead, cheaper and often better for planning, at the cost of outputs a human cannot eyeball. Either way the data has to pair observations with the actions that produced them; passive video alone teaches what tends to happen, not what your intervention causes. Evaluation is the unsolved part: a sharp-looking prediction can still be physically wrong, so teams score world models on downstream control tasks, not on how convincing the imagined footage looks.

What does a world model change for business software?

Nothing this quarter, and possibly a lot later. Today the practical descendants are vision language models that ground software in what a screen or camera actually shows, and simulation-style testing such as agent simulation that rehearses an agent before production. Treat world models as a research direction to track, not a product category to buy.

Frequently asked questions

What is the difference between an LLM and a world model?

An LLM predicts the next token in a sequence of text, so it models descriptions of the world. A world model predicts the next state of an environment, so it models the world itself: objects, physics, and what an action changes. The two are complementary rather than competing.

How do you build a world model?

By training on observation data such as video, simulation or robot sensor streams, so the model learns environment dynamics rather than language. Approaches differ by lab, from generative video models to embedding-prediction architectures, and the field is still far from a standard recipe.

Summarize this page with

Train your team to build this