World Labs debuts Atlas, an omni world model, in early access
TL;DR
- Atlas is pretrained from scratch on text, images, video, and 3D natively, a structural difference from video models with bolt-on reconstruction.
- On sparse-view 3D reconstruction, Atlas scores 25.3 mean absolute-relative pointmap error versus 28.7 for the next-best specialist model.
- Human raters preferred Atlas over rival video models in 75 to 94 percent of trials, varying by competitor tested.
World Labs is releasing Atlas, a single model it says handles text, images, video and 3D through one "shared spatial context," and it is publishing head-to-head numbers to back the claim. In a company blog post, the Fei-Fei Li cofounded startup describes Atlas as "a multimodal autoregressive diffusion transformer," pretrained from scratch, with camera geometry treated as a native input rather than something coaxed out of a text prompt.
On camera-controlled generation, the post reports third-party human raters picked Atlas over MiniMax H3 75% of the time, over Gemini Omni Flash 81%, over Happy Horse 1.1 86%, over FLUX 3 93%, and over Seedance 2.5 94%. The showpiece demo is "a 1 minute video at 1440p resolution" generated from a small number of reference images along a hand-designed camera path.
For 3D reconstruction, Atlas takes "as few as two or three images" and posts mean absolute-relative pointmap errors of 8.6 on DTU, 9.3 on ETH3D and 12.4 on ScanNet, lower than every reproduced open-source specialist across seven benchmarks in the results table. World Labs frames it flatly: "Despite its generality, Atlas outperforms the best specialized open-source reconstruction models."
The write-up is upfront about a framing risk in the video comparison. Baselines "do not accept cameras as a native input format," so their camera paths were described in text prompts, and the post concedes "more sophisticated prompt engineering or creative multimodal prompts could improve camera following for some models."
Robotics is a pitched application. World Labs says it captured "two large environments with a cell phone video, using 24 frames each" for reconstruction, then simulated different robots navigating the scenes and rendered what body-mounted cameras would see. The same trick underlies a "bullet time" reframing demo, filmed, per the post, "using just a few cell phones and action cameras."
Atlas is "entering early access with select partners" and will "power future versions of Marble and other products from World Labs." No pricing, latency, compute footprint, or partner names appear in the post. It lands the same day Runway announced Solaris, an "interface world model" aimed at live UIs, one of several world-model launches our multimodal feed logged today.
What others are reporting
-
SiliconANGLE Read →
Emphasizes the robotics application: developers photograph physical spaces with a smartphone and reconstruct them as 3D simulations for safe robot training.
Atlas is a breakthrough multimodal world model that aims to bridge the gap between simulated environments and physical reasoning.
-
CryptoBriefing Read →
Traces the lineage from Marble (Nov 2025) through World API (Jan 2026) and the SceniX acquisition (Jul 2026), framing Atlas as a multi-year architecture build.
Atlas can produce pixel-perfect, camera-controlled imagery and video at resolutions up to 1440p, running for durations of up to one minute.
-
AlphaSignal Read →
Leads with concrete benchmark numbers across DTU and T&T datasets and notes 75 to 94 percent human preference over rivals in camera control trials.
Atlas is pretrained from scratch to natively operate on text, images, video, and 3D as a multimodal autoregressive diffusion transformer.
-
Radiance Fields Read →
3D-specialist outlet details cofounder Ben Mildenhall's NeRF origins and explains how Atlas writes scenes as point clouds or Gaussian splats as native output.
Atlas natively operates on both 2D image frames and 3D depth maps, which lets it write worlds out as point clouds or as splats.
-
Superpower Daily Read →
Flags that pricing, partner names, and public availability remain undisclosed, and that production validation by early-access partners is still ahead.
Atlas accepts one to six reference images and a manually designed camera path for up to one minute of 1440p video.
Shared on Bluesky by 2 AI experts
-
World Labs introduces Atlas. A world model that constructs 3D models from images and simulates physics to some degree. www.worldlabs.ai/blog/atlas www.youtube.com/watch?v=hzvX...
View on Bluesky →
Originally reported by worldlabs.ai
Read the original article →Original headline: Fei-Fei Li's World Labs Unveils Atlas, a Multimodal World Model Generating 1-Minute 1440p Video