Realtime AI News
Fei-Fei Li's World Labs unveils Atlas, billed as the world's first multimodal world model
World Labs, founded by Fei-Fei Li, has released Atlas, which it calls the world's first multimodal world model — generating images and video with pixel-level camera control from a single photo and reconstructing 3D scenes. The model can also turn a few photos into realistic robot training data, a key step on the real-to-sim path for embodied AI.
World Labs, the company founded by AI pioneer Fei-Fei Li, has released Atlas, which it calls the world's first multimodal world model. According to a report by Chinese outlet QbitAI, the model goes beyond generating interactive video clips — its stated mission is to "model the world, move the camera, simulate space and time," producing images and video frames with pixel-level camera control while performing 3D reconstruction.
From a single image, Atlas can generate images and video with pixel-level camera control, at up to one minute of 1440p footage. Given anywhere from one to dozens of input images, it can reconstruct real-world scenes, producing both novel-view image frames and explicit 3D representations that reportedly outperform state-of-the-art models trained specifically for 3D reconstruction. It can also generate images and 360-degree panoramas from text, following complex prompts and rendering text accurately.
For robotics, the model's spatiotemporal simulation is the standout capability: Atlas can model both space and time from an input video, change the viewing angle of existing footage, and support real-to-sim workflows. Feeding it a few photos yields realistic RGB and depth data, allowing robots to train and be tested in many more simulated environments.
Technically, Atlas is an "omni model" built on a newly designed multimodal autoregressive diffusion Transformer architecture. Text, images, video and 3D data are anchored in three-dimensional space to form a spatial context, from which the model generates multimodal output. As a Rectified Flow diffusion model, it can balance generation speed and quality by adjusting the number of denoising steps.
In quantitative evaluations on camera-controlled generation and 3D reconstruction, Atlas beat state-of-the-art video models on camera control — with the advantage growing as camera trajectories become more complex — and outperformed the best open-source 3D reconstruction models on sparse-view reconstruction. The team says Atlas was designed for scaling from the start, with early evidence that its capabilities keep improving as it grows.
Atlas ties together Fei-Fei Li's world-model agenda this year: in June she classified world models into renderers, simulators and planners, arguing the simulator matters most; in July World Labs acquired robot-simulation startup SceniX and published its Real-to-Sim-to-Real system, identifying the lack of cheap, controllable training experience as robotics' biggest bottleneck. Atlas now completes the path from real photos and video to 3D space to robot sensor views. NVIDIA's head of robotics, Jim Fan, publicly called it a big step for real-to-sim in robotics.
Atlas is currently in early access for select partners, with applications open on the World Labs website. The company says Atlas will become the foundation model for future versions of Marble and other World Labs products, hinting at a general world-model platform that can simulate space and time.
Why it matters
World models are moving from generating video to simulating an interactive 3D world. By linking generative AI directly to robot training, Atlas could sharply lower the cost of embodied-AI training data and reshape workflows in visual effects, autonomous driving and robotics.
Nearby Updates
All09/02, 09:31
Anthropic admits security failures behind AI hacking incidents, says models 'not perfectly aligned' with human values
Anthropic has admitted that a series of hacking incidents involving its models reflected a "failure of operational security" and said it has tightened its testing procedures. In a new blog post, the company acknowledged its technology is "not perfectly aligned" with human values and goals, citing motivated reasoning and recklessness as two alignment failures found in the incidents.
09/02, 08:13
Zhaogang.com-W launches NextB2B, an AI agent for the trading industry
Zhaogang.com-W (06676) has launched NextB2B, an AI agent built for the trading industry, according to a Moomoo report. The move extends the company from its platform business toward AI-powered tools for trade workflows.
09/02, 06:50
ClawBench benchmark: Claude, GPT and Gemini agents fail 33% of tasks on real websites
A new benchmark called ClawBench tests AI agents on real websites, and the results show Claude, GPT and Gemini agents collectively failing 33% of tasks. The findings expose the gap between frontier agents' strong lab scores and their reliability in messy real-world conditions.
09/02, 06:08
AfterQuery reportedly becomes Y Combinator's fastest-ever unicorn at $3.2B valuation
TechCrunch reports that AfterQuery, an AI model-training startup, has raised a round valuing it at $3.2 billion, reportedly making it Y Combinator's fastest-ever unicorn. The valuation comes just five months after the company announced its $30 million Series A at a $300 million valuation in April.