Atlas: A World Model for Spatial Intelligence
Points and comments are a snapshot, not live.
World Labs unveils Atlas, an omni world model for 3D spatial intelligence from sparse images.
Atlas is a multimodal autoregressive diffusion transformer pretrained on text, images, video, and 3D. It performs camera-controlled generation (up to 1 min video at 1440p), spatial reconstruction from one to dozens of input images (outperforming specialized 3D models), space-time simulation for robotics, and text-to-image generation including 360 panoramas. The model uses a spatial context where each image is grounded at a 3D position, enabling pixel-perfect camera control and consistent world generation. Atlas can output both 2D frames and explicit 3D representations (point clouds, 3D Gaussian splats). It will power World Labs' Marble product and other applications.
What commenters are saying
Commenters are impressed by Atlas's ability to reconstruct 3D spaces from sparse inputs, with potential for robotics and creative workflows. A World Labs cofounder and the project lead answered questions: Atlas maintains 3D consistency as the camera moves, uses explicit camera poses as native input (unlike models like Genie 3 that use raw keyboard commands), and can manage context creatively via "context juggling." Some noted the model appears best at frozen-time reconstructions with limited scene motion, though it can handle some movement. The term "world model" drew debate over its overloading. Concerns about hallucination in unknown areas were addressed: Atlas can operate in faithful reconstruction mode or imagine completions depending on the application.