Why are AI labs building world models?
Because next-token prediction runs out of world. An LLM has read about gravity but never watched anything fall, which caps how well it can plan physical or spatial work. A world model gives an agent a place to rehearse: it can simulate an action internally, check the predicted outcome and only then act. That matters for robotics, autonomous operations and any computer use task where a wrong click has consequences, and it is why several frontier labs now describe world models as the next capability jump after reasoning models.
How does a world model actually work?
The core loop is encode, predict, compare. The model compresses each observation, a video frame or a sensor reading, into an internal state, then predicts how that state changes when an action is taken. Training pushes the prediction toward what actually happened, so over millions of transitions the model absorbs the dynamics of its environment: what moves, what collides, what an intervention changes. Prediction conditioned on actions is the defining ingredient, because it turns the model into something an agent can query. Try action A in imagination, read the predicted outcome, compare it with action B, then commit. The imagined attempt costs almost nothing; the real one, in physical AI especially, does not.
How is a world model trained and evaluated?
Two families dominate. Generative approaches predict the next observation directly, pixels included, which makes results easy to inspect but spends most of the model’s capacity on visual detail that rarely matters for a decision. Latent approaches predict in the compressed state space instead, cheaper and often better for planning, at the cost of outputs a human cannot eyeball. Either way the data has to pair observations with the actions that produced them; passive video alone teaches what tends to happen, not what your intervention causes. Evaluation is the unsolved part: a sharp-looking prediction can still be physically wrong, so teams score world models on downstream control tasks, not on how convincing the imagined footage looks.
What does a world model change for business software?
Nothing this quarter, and possibly a lot later. Today the practical descendants are vision language models that ground software in what a screen or camera actually shows, and simulation-style testing such as agent simulation that rehearses an agent before production. Treat world models as a research direction to track, not a product category to buy.