📎 Webclip
Zhuokai Zhao on X: Five categories of world models
The post says that almost nobody means the same thing by “world models” and groups the term into five categories. It contrasts JEPA’s latent prediction, spatial intelligence’s 3D structure, learned simulation through generated experience, NVIDIA’s infrastructure layer, and active inference as a different theory of intelligence.
Reading notes#
- AMI Labs and World Labs are both betting on world models, but the post argues the term covers different approaches.
- JEPA predicts in learned latent space instead of reconstructing pixels, which the author presents as better for physical understanding.
- V-JEPA 2 is described as a proof point because a small amount of robot data was enough to support zero-shot planning.
- Spatial intelligence is framed as building persistent 3D environments with geometry, depth, and viewpoint changes.
- Marble is described as generating 3D scenes from images, text, video, or 3D layouts, with orbiting, editing, and mesh export.
- Learned simulation combines generative video models and latent-space RL, both focused on simulating how actions change environments over time.
- Genie 3 is presented as a navigable environment generated from text, while Dreamer 4 is described as reaching Minecraft diamonds through imagination.
- NVIDIA Cosmos is presented as a platform for world models, covering data pipelines, tokenization, training, deployment, and model families for prediction, transfer, and reasoning.
- Active inference is presented as a separate framework from deep learning, based on minimizing surprise through Bayesian updating.
- AXIOM is described as using object-centric structured models and hierarchical multi-agent control with online inference.
- The post concludes that these categories are not really competitors because they solve different subproblems.
