Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

UniWorld-View Tops Fei-Fei Li Team's World Model Leaderboard, Fully Open-Sourced with Ascend NPU Support

UniWorld-View, developed by Toozhan AI in collaboration with Peking University and Peng Cheng Laboratory, has topped the world model leaderboard maintained by Fei-Fei Li's team. The model supports novel view synthesis from a single image or monocular video, is compatible with domestic Ascend NPUs, and has released all code and weights as open source.

Published

The world model leaderboard from Fei-Fei Li's team has a new top spot. UniWorld-View, jointly developed by Toozhan AI (兔展智能), Peking University, and Peng Cheng Laboratory, has claimed the top position. The model is already compatible with domestic Ascend NPUs and has been fully open-sourced with all code and weights available on GitHub.

UniWorld-View specializes in Novel View Synthesis (NVS), enabling AI to hallucinate unseen camera perspectives from a single image or monocular video. Users can capture a casual video or even just a single photo, and the model generates new-angle videos with freely controllable camera trajectories — push, pull, pan, tilt, or even a full 360-degree orbit around the scene, with precise camera pose control.

Traditional approaches such as NeRF and 3D Gaussian Splatting rely on the reconstruction paradigm, requiring multi-view inputs. When faced with monocular videos that have minimal coverage, these methods suffer from hollow artifacts or outright failure. Most generative approaches, meanwhile, simply feed camera poses as abstract condition vectors into diffusion models, lacking explicit 3D geometry — causing viewpoint drift and loss of control at larger baselines.

UniWorld-View's core technical innovation is its occlusion-aware point cloud rendering. The team identified a flaw in prior methods: conventional point clouds are collections of isolated points lacking connectivity or surface orientation cues. During large-baseline view changes, foreground textures get stretched across backgrounds like gum, or background pixels bleed through foreground objects. UniWorld-View uses Double-Reprojection to identify occluded regions and remove erroneous signals, and estimates normal vectors for each point — culling back-facing points from rendering entirely — providing clean geometric conditions for the diffusion model.

During generation, UniWorld-View employs a dual-stream conditioning strategy: one stream receives the point cloud render for positional accuracy, the other receives the source video for complete content. This ensures output aligns with the target viewpoint while preserving the original scene appearance. A further milestone is its end-to-end pipeline from monocular video to multi-view video to full 4D scene — by generating multiple novel views and feeding them into 4D Gaussian Splatting optimization, the model produces freely navigable 4D scenes from casual monocular capture.

Toozhan AI was founded by Peking University alumni and young leading researchers in computer vision from PKU. The company has developed the Open-Sora Plan and UniWorld series of visual AI models. Open-Sora Plan was the industry's first open-source text-to-video model and led global code citations in 2024, while the UniWorld series pioneered unified understanding-and-generation architectures. Toozhan previously released UniWorld-V2 and UniWorld-V2.5, matching GPT-Image-2's generation quality on infographics and text-dense scenarios.

The training data for UniWorld-View comes from YuanKong Intelligence, a Peking University-incubated startup. Importantly, the model's compatibility with domestic Ascend NPUs signals continued progress in China's domestic AI hardware infrastructure for large-scale visual model training and inference. Toozhan AI also disclosed plans to launch RabbitVis, a design-production tool integrating model capabilities, aiming to bring visual AI into commercial design workflows.

As competition in the world model space intensifies, UniWorld-View's combination of a top leaderboard ranking and fully open-sourced stack positions it to provide foundational capabilities for embodied intelligence, autonomous driving, and cinematic production — applications that demand rich 3D scene understanding.

Why it matters

UniWorld-View topping the leaderboard with full open-source release marks a significant breakthrough for Chinese visual AI in the world model space, and its Ascend NPU compatibility strengthens the domestic AI infrastructure ecosystem for embodied intelligence and 4D scene generation.

UniWorld-View世界模型兔展智能北京大学鹏城实验室昇腾
Back to realtime news

Nearby Updates

All

07/24, 22:42

Hefei-Backed AI Unicorn HiDream.ai Raises $290M in Three Months, Signals Multimodal Shift

HiDream.ai, a multimodal AI company based in Hefei, announced a 1.5 billion yuan ($207M) Series C round, completing over 2.1 billion yuan ($290M) in three consecutive funding rounds within three months. The company is now valued at over $1 billion, joining the global AI unicorn club with backing from national pension funds, state capital, and film industry investors.

07/24, 21:46

Xiaomi new phone passes certification with a dozen AI models including DeepSeek, ERNIE Bot, and Tongyi Qianwen

A Xiaomi phone model 2608BPX34C has passed network access certification in China, filing support for a dozen large language models including DeepSeek, ERNIE Bot, Tongyi Qianwen, Zhipu AI, and Xiaomi's own models. Industry watchers speculate the filings are preparation for the upcoming HyperOS 4 system.

07/24, 21:36

OpenAI brings ChatGPT Voice to the desktop app, enabling hands-free agent control

OpenAI on Thursday updated its ChatGPT desktop app with voice mode support, letting users speak commands to control AI agents and perform computer tasks. The feature uses the ChatGPT-Live voice model family launched earlier this month and works with ChatGPT Work and Codex for multi-step operations.

07/24, 23:09

Midjourney Acquires Astrology App Co-Star, Bringing Two-Dozen Employees Onboard

AI lab Midjourney has acquired the popular astrology app Co-Star, bringing its team of roughly two dozen employees onboard. The acquisition signals Midjourney's ambition to expand beyond its Discord-based image generation roots and potentially build its first standalone app.