Mainstay Digital
Start a request

Research

World Labs Launches Atlas, a Spatial World Model That Generates Geometrically Consistent 3D Video from Single Images

Fei-Fei Li's startup ships a multimodal autoregressive diffusion transformer that produces 1440p video with 3D Gaussian splat outputs, moving the field from text-prompted frame sequences toward geometry-native spatial generation.

M

Mainstay Communications Team · · 6 min read

In brief

World Labs unveiled Atlas on September 1, 2026, an omni world model that generates up to one minute of 1440p camera-controlled video from a single image and reconstructs scenes as point clouds or 3D Gaussian splats. Human raters preferred Atlas over Google's Gemini Omni Flash in 81% of camera-controlled generation trials, and Atlas posted a mean absolute-relative pointmap error of 25.3 against 28.7 for the next-best specialist model. The launch arrives fourteen months after World Labs' first product, Marble, and accelerates the company's $1.23 billion bet that spatial intelligence becomes the organizing layer beneath generative AI.

What Atlas Does

World Labs describes Atlas as a multimodal autoregressive diffusion transformer pretrained from scratch on text, images, video, and 3D data. Each input anchors to a 3D position inside a shared spatial context rather than being processed as a flat sequence, which lets the model hold geometric consistency across views it was never shown.

From one to six reference images and a manually designed camera path, Atlas generates as much as one minute of video at 1440p. It can also rebuild a room from as few as two or three ordinary phone photos, outputting the result as a point cloud or a 3D Gaussian splat file ready for downstream pipelines.

The architecture combines the piece-by-piece generation of a language model with the diffusion-based quality of video models. World Labs calls the result rectified-flow latent diffusion with KV caching, a combination the company says lets Atlas scale with compute in a predictable way.

How It Differs from Prior Video Models

The central distinction is how camera control works. Models like Google's Gemini Omni Flash accept text descriptions of movement ("pan right," "crane down") and interpret them as prompts. Atlas accepts camera trajectories and geometry as direct numerical inputs. The shot is specified in the same coordinate space the model reasons in.

That matters because text-prompted camera moves are approximations. A model asked to "pan right" learns a statistical average of what that phrase implied during training. A model handed a trajectory and depth map has the constraint locked in before it generates a single frame.

Prior generations of video models worked on time-ordered pixel sequences. Atlas grounds generation in 3D coordinates first and renders views of that space second. The company's launch post describes the goal as a model that "models the world, moves the camera, and simulates space and time."

The Benchmark Picture

World Labs measured 3D reconstruction error using mean absolute-relative pointmap error, averaged across seven standard datasets. Atlas scored 25.3; Pi3X scored 28.7; Pi-cubed scored 34.7; VGGT-Omega 1B scored 36.4; Depth Anything 3 scored 39.3; MapAnything scored 47.7. Lower is better, and the scores are reported in units of 10 to the negative 3.

On camera-controlled generation, human raters preferred Atlas over Seedance 2.5 in 94% of trials, over FLUX 3 in 93%, over Happy Horse 1.1 in 86%, over Gemini Omni Flash in 81%, and over MiniMax H3 in 75%.

One caveat World Labs acknowledges in its own materials: Atlas received native camera trajectories during evaluation while competitors received text descriptions. The company says more sophisticated prompting could improve competitor results, which makes the absolute preference percentages less meaningful than the relative ordering.

Robotics as the First Market

World Labs has positioned robotics as the clearest near-term use case. Atlas can reconstruct a physical room and generate the RGB images and depth readings a simulated robot's cameras would see along any path through it. Developers can then swap out objects, lighting, or backgrounds to generate diverse training data without staging new real-world captures for every variation.

In July 2026, World Labs acquired SceniX, a robotics company co-founded by Yunzhu Li and Changxi Zheng that builds high-fidelity simulation platforms for training robots. World Labs said in August 2026 that its real-to-sim-to-real engine, incorporating SceniX's work, had run five robot platforms for an hour each without human intervention in a controlled test.

Atlas feeds into that loop by making the 3D reconstruction from casual video fast enough to be practical. A warehouse floor, a manufacturing cell, or an assembly station can become a training environment with a phone and an Atlas API call.

The Company Behind It

World Labs was founded in January 2024 by Fei-Fei Li, Justin Johnson, Ben Mildenhall, and Christoph Lassner. Li built ImageNet and ran AI at Google Cloud. Mildenhall co-created NeRF (Neural Radiance Fields) at Google Research. Lassner developed Pulsar, a sphere-based differentiable renderer that laid groundwork for Gaussian splatting, during his time at Meta Reality Labs and Epic Games.

The company raised $230 million at its September 2024 launch, then closed a $1 billion Series B on February 18, 2026, with AMD, Autodesk, Emerson Collective, Fidelity Management and Research, NVIDIA, and Sea Group participating. Autodesk anchored the round with a $200 million investment. Total capital raised stands at $1.23 billion; Bloomberg reported the round valued World Labs at $5 billion, a figure the company declined to confirm publicly.

World Labs launched Marble in November 2025 as its first commercial product. Marble takes a single image or text prompt and outputs a navigable 3D world at subscription tiers ranging from free to $95 per month. Atlas is positioned as the underlying model that will power future Marble iterations.

Access and What's Missing

Atlas entered early access on September 1 with an unnamed set of select partners. Applications go through a typeform form on World Labs' website; no general availability date is set and no pricing is published.

The launch arrived without an academic paper, a model card, a parameter count, or any named early-access partners. World Labs describes its training corpus only as "a large diverse multimodal dataset." Independent verification of the benchmark results hasn't been published as of this writing. Atlas used VGGT-Omega 1B, one of its comparison baselines, which had carried an August 18, 2026 benchmark-contamination warning that its published results might be inflated. The VGGT-Omega team withdrew that warning on September 7, 2026, releasing training code, a retrained checkpoint, and a reproduction report that found the corrected results broadly comparable to the originals. World Labs still hasn't identified which VGGT-Omega checkpoint it benchmarked against.

That transparency gap is worth watching. The benchmark numbers are plausible, the architectural claims cohere with published research on spatial transformers and Gaussian splatting, and the SceniX acquisition shows a real robotics path. Production validation depends on what the early-access partners report.

Sources

Companies mentioned

Mainstay Digital · Spatial and 3D visualization for industrial companies