HeadlinesBriefing favicon HeadlinesBriefing.com

Atlas: World Labs' New Spatial Intelligence Model

Hacker News •
×

World Labs introduces Atlas, a next-generation world model designed for spatial intelligence. Atlas is an omni model pretrained from scratch to natively operate on text, images, video, and 3D. It is a multimodal autoregressive diffusion transformer that combines all inputs into a shared spatial context to generate what comes next while staying consistent in 3D and imagining beyond observed data.

Atlas scales with increased training compute and supports world generation, reconstruction, and simulation. Key capabilities include Camera-Controlled Generation, producing up to 1 minute of 1440p video from one or more images with pixel-perfect camera control; Spatial Reconstruction, generating novel views and explicit 3D outputs from input images, outperforming state-of-the-art 3D models; Space-Time Simulation, enabling Real-to-Sim workflows for robotics by modeling space and time from videos; and Image Generation, creating images and 360 panoramas from text with complex prompt adherence and diverse visual styles. Atlas will power future versions of Marble and other World Labs products.

The model uses precise camera geometry as a native input, allowing exact shot framing and motion control. By grounding each image at a 3D position in space, Atlas forms a spatial context that enables creative control, such as placing unrelated images in 3D space to generate smooth interpolations imagining doorways, hallways, and transitions. For long videos, Atlas combines camera movement and spatial context management, putting users in the director's chair for full scene staging.