a16z|Sep 04, 2026 14:34
World Labs co-founders Fei-Fei Li, Justin Johnson, Ben Mildenhall, and a16z's Martin Casado on Atlas, a world model for spatial intelligence:
LLMs are built on next token prediction. Video models are built on next frame prediction. Atlas is built on new view prediction, and it's the first model to unify pixel generation and pixel reconstruction, two problems computer vision has kept in separate tracks for half a century.
The practical result is a 50 to 100x reduction in what it takes to digitally capture a 3D representation of a space. Previously, you needed 100 to 300 photos of a single room. Atlas can work from just three.
In this conversation, they get into the slow motion shot from The Matrix that took hundreds of cameras and now takes three iPhones, the overnight Slack message that made them bet the company in five seconds, why robotics is bottlenecked on data rather than chips, and the case that new view prediction is AI-complete.
00:00 Intro
01:50 The Matrix slow motion scene now takes three iPhones
02:48 Why new view prediction is the primitive
07:10 Unifying generation and reconstruction
11:15 Gaussian splats became the bottleneck
14:17 Dense capture used to mean 300 photos
17:30 Why reconstruction needs generation to fill the gaps
18:44 The LLM lesson image models missed
23:39 The video that made them go all in
28:04 3D design is 95% revisions
30:50 The problem in robotics is data, not chips
32:48 Why a robot policy can't be trained like an image model
34:44 When the simulator becomes the planner
36:45 Frozen time required footage full of movement
40:57 Why new view prediction is AI-complete
42:43 Nature gave animals eyes but not trees
YouTube: https://www.youtube.com/watch?v=qn1QDDBnTA0
@drfeifei @jcjohnss @BenMildenhall @theworldlabs @martin_casado(a16z)
Share To
HotFlash
APP
X
Telegram
CopyLink