On October 6, TechCrunch profiled Mirror Particle, a startup building a world model of human behavior, beside three others that raised a combined $768 million for the same idea inside a year. The label is suddenly everywhere, and it is worth asking what it buys you. A world model predicts the next state of an environment rather than the next word, and three of the four flagship systems, Marble, Cosmos 3 and V-JEPA 2, ship today. The fourth, DeepMind’s Genie 3, is the one everyone has seen and nobody outside a research preview can touch. This post separates the two things the term now means, walks all four systems with what each licenses and exports, and says where one actually fits in a production stack.
What’s the difference between LLM and World Model?
An LLM is trained on text to predict the next token; a world model is trained on video, sensor and action data to predict the next state of the world. Google DeepMind’s Genie 3 announcement defines world models as “AI systems that can use their understanding of the world to simulate aspects of it, enabling agents to predict both how an environment will evolve and how their actions will affect it”. The practical difference is the output: an LLM hands you language, a world model hands you a scene, a frame or a prediction you can act against.
The term covers two different architectures, a distinction Ben Dickson’s TechTalks analysis draws cleanly. The first is generative: systems like Marble and Genie 3 create and simulate an external environment you can walk through. The second is predictive: an internal system an agent uses to anticipate what happens next, the way Meta’s V-JEPA 2 predicts the outcome of a robot’s action without rendering a single pixel. Vendors use one word for both, and the two solve different problems, so every evaluation should start by asking which kind is on the table.
Genie 3 is the demo you cannot build on
Google DeepMind announced Genie 3, in the post quoted above, on August 5, 2025 as a general purpose world model that generates interactive environments from a text prompt, navigable in real time at 24 frames per second and 720p, staying consistent for a few minutes. Consistency here is emergent rather than built on an explicit 3D representation, and the model supports what DeepMind calls promptable world events, text commands that change the weather or add objects mid-session.
The catch sits in the same announcement: Genie 3 launched as a limited research preview for a small cohort of academics and creators, and the announcement itself lists limited interaction duration, a few minutes rather than extended hours, among its constraints. That status has outlasted the news cycle; DeepMind’s page still describes early access the same way today. DeepMind’s stated destination is training agents in an unlimited curriculum of generated environments. For a builder with a 2026 roadmap, Genie 3 is a capability forecast, not a dependency.
Marble ships real assets today, from a company mid-acquisition
World Labs made Marble generally available on November 12, 2025, and it remains the system with the clearest production story. Marble creates 3D worlds from text, images, video, or coarse 3D layouts, and the generated worlds can be “exported as Gaussian splats, meshes, or videos”. Export is the load-bearing word: the output is an asset in your pipeline, not a session inside someone else’s viewer. The release includes Chisel, an editor for sculpting generated worlds directly in 3D, and Spark, an open-source renderer for displaying Gaussian splats in the browser on top of THREE.js. If your problem looks like environment creation, a game level, a training scene for a digital twin, a navigable space from a site video, this is the one you can buy today.
One diligence note belongs in any Marble evaluation: World Labs signed a definitive agreement to join AMD on September 28, 2026, with Fei-Fei Li joining as Executive Vice President and Chief Scientist and the transaction expected to close by the end of 2026. Nothing about the product changed at announcement, but a vendor mid-acquisition is a roadmap risk you price in. The open-weight hedge exists: Tencent’s HunyuanWorld 1.0 generates explorable 3D scenes from words or pixels under a community license, downloadable from Hugging Face.
Cosmos 3 is fully open and past 10 million downloads
NVIDIA took the opposite route from both. Cosmos launched at CES on January 6, 2025 as a platform of world foundation models for physical AI, with the launch announcement naming 1X, Agility, Figure AI, Uber, Waabi and XPENG among the first adopters and claiming its pipeline can process and label 20 million hours of video in 14 days on Blackwell hardware. Jensen Huang’s framing from that release is the sector’s thesis in one line: “The ChatGPT moment for robotics is coming.”
Eighteen months on, the numbers back the open strategy. NVIDIA’s developer announcement of July 27, 2026 reports Cosmos models passing 10 million downloads on Hugging Face, with Cosmos 3 Edge bringing an open 4B world model to edge devices. The current generation, Cosmos 3, is an omni-model generating across text, image, video, sound and action on a Mixture of Transformers architecture, released under the Linux Foundation’s OpenMDW1.1 license with post-training scripts on GitHub for each modality. For NVIDIA the model is the demand generator for the hardware, which is exactly why a builder gets weights, fine-tuning recipes and synthetic data tooling instead of a waitlist.
V-JEPA 2 predicts outcomes instead of rendering pixels
Meta released V-JEPA 2 on June 11, 2025, and it is the clearest shipped example of the second, predictive meaning of the term. Trained on video, it gives agents three capabilities Meta names explicitly: understanding, predicting and planning. It does not generate a watchable world. It predicts how a scene will respond to an action, which was enough for robots in Meta’s labs to perform reaching, picking up an object and placing it in a new location, on a model you can download: the V-JEPA 2 checkpoints on Hugging Face carry an MIT license and the ViT-L encoder alone logged 277,854 downloads in the month we checked.
For a production team this is the quiet option. A vision-language model answers questions about a frame; V-JEPA 2 is an encoder that anticipates the next one, a building block for manipulation, navigation and video understanding rather than a creative tool. Dickson’s analysis above closes on the synthesis worth remembering: generative worlds like Marble’s to train in, efficient predictive models like V-JEPA as the agent’s internal physics.
The money arrived before the products
The funding curve says this category will be pitched to you long before it stabilizes. Yann LeCun’s AMI Labs raised $1.03 billion at a $3.5 billion pre-money valuation in March 2026 after reportedly seeking 500 million euros and closing around 890 million, backed by NVIDIA, Samsung and Toyota Ventures among others, and TechCrunch noted World Labs had secured $1 billion the month before. The behavioral wing is just as hot: the Mirror Particle profile from the intro counts Simile at $200 million on a $2 billion valuation, Aaru at $88 million on $1 billion, and Humans& announcing a $480 million seed at a $4.48 billion valuation in January.
AMI Labs CEO Alexandre LeBrun called the branding wave himself, telling TechCrunch in the same funding report: “In six months, every company will call itself a world model to raise funding.” Mirror Particle’s CEO Abhivyakti Ahuja pitches the same anti-LLM thesis more vividly, saying that using an LLM to model human behavior is “like bringing a super soaker to Niagara Falls”. Both statements serve the people raising the money, which is precisely why the label on the deck tells you nothing. A frontier model claim is checkable; “world model” in a pitch currently is not.
The trap: evaluating the category on its flashiest demo. The demo is Genie 3, a research preview; the licenses you can ship on this quarter are Marble’s, Cosmos 3’s and V-JEPA 2’s, and they solve three different problems.
Where a world model fits in your stack this quarter
Interest is measurably rising: the keyword research we run for our own AI glossary puts “world model” at 6,600 monthly searches in the US and 2,460 across Poland, the UK, Germany and the Netherlands, with a rising 12-month trend, so expect the question from your own stakeholders soon. Our answer, for the teams we build with: if your product manipulates documents, tickets and systems of record, you do not need a world model, and the agents ranked in the most used AI models remain the shipping path. A world model enters the stack when the problem is physical or spatial. Choosing an environment generator, start with Marble and price the AMD close into the contract. Training robots or vehicles, start with Cosmos 3, because open weights plus post-training scripts beat a demo you cannot license, the same logic that made the harness, not the model, the asset you own. Building perception for manipulation, start with V-JEPA 2 under MIT. And if the work is on a factory floor, the robot that pays first is usually the boring one, as the sanding cell in our manufacturing piece showed.
The takeaway
A world model predicts states, an LLM predicts tokens, and the category currently spans two architectures sold under one name: generative worlds you walk through and predictive models an agent plans with. Today the buildable list is short and concrete. Marble is generally available and exports Gaussian splats, meshes and video; Cosmos 3 is fully open under OpenMDW with 10 million downloads behind it; V-JEPA 2 is MIT-licensed and already moves lab robots. Genie 3, the most impressive demo of the four, remains a research preview. Evaluate the license and the export path before the demo, and treat “world model” on a pitch deck the way you treat any label that predates its product.
Sources
- Genie 3: a new frontier for world models, Google DeepMind
- Marble, a multimodal world model, World Labs
- World Labs is joining AMD, World Labs
- NVIDIA Cosmos, NVIDIA
- NVIDIA launches Cosmos world foundation model platform, NVIDIA Newsroom
- 10M downloads and counting: NVIDIA Cosmos, NVIDIA Developer Forums
- Our new model helps AI think before it acts, Meta Newsroom
- V-JEPA 2 model card, Meta on Hugging Face
- HunyuanWorld 1.0 model card, Tencent on Hugging Face
- Mirror Particle is building a world model of human behavior, TechCrunch
- Yann LeCun’s AMI Labs raises $1.03 billion to build world models, TechCrunch
- What to know about World Labs’ Marble generative world model, TechTalks