worldlabs.ai via Hacker News

World Labs debuts Atlas, an omni world model, in early access

TL;DR

  • Atlas is a multimodal autoregressive diffusion transformer that generates up to one minute of camera-controlled 1440p video from a small number of reference images.
  • On sparse-view 3D reconstruction, Atlas beats every reproduced open-source specialist across DTU, ETH3D, ScanNet and other benchmarks per World Labs' own numbers.
  • Atlas enters early access with unnamed partners and will power future versions of Marble and other World Labs products.

World Labs is releasing Atlas, a single model it says handles text, images, video and 3D through one "shared spatial context," and it is publishing head-to-head numbers to back the claim. In a company blog post, the Fei-Fei Li cofounded startup describes Atlas as "a multimodal autoregressive diffusion transformer," pretrained from scratch, with camera geometry treated as a native input rather than something coaxed out of a text prompt.

On camera-controlled generation, the post reports third-party human raters picked Atlas over MiniMax H3 75% of the time, over Gemini Omni Flash 81%, over Happy Horse 1.1 86%, over FLUX 3 93%, and over Seedance 2.5 94%. The showpiece demo is "a 1 minute video at 1440p resolution" generated from a small number of reference images along a hand-designed camera path.

For 3D reconstruction, Atlas takes "as few as two or three images" and posts mean absolute-relative pointmap errors of 8.6 on DTU, 9.3 on ETH3D and 12.4 on ScanNet, lower than every reproduced open-source specialist across seven benchmarks in the results table. World Labs frames it flatly: "Despite its generality, Atlas outperforms the best specialized open-source reconstruction models."

The write-up is upfront about a framing risk in the video comparison. Baselines "do not accept cameras as a native input format," so their camera paths were described in text prompts, and the post concedes "more sophisticated prompt engineering or creative multimodal prompts could improve camera following for some models."

Robotics is a pitched application. World Labs says it captured "two large environments with a cell phone video, using 24 frames each" for reconstruction, then simulated different robots navigating the scenes and rendered what body-mounted cameras would see. The same trick underlies a "bullet time" reframing demo, filmed, per the post, "using just a few cell phones and action cameras."

Atlas is "entering early access with select partners" and will "power future versions of Marble and other products from World Labs." No pricing, latency, compute footprint, or partner names appear in the post. It lands the same day Runway announced Solaris, an "interface world model" aimed at live UIs, one of several world-model launches our multimodal feed logged today.