Realtime AI News
HiDream.ai releases HiDream-O1-Video-1.0, a natively omni-modal video model that lands in the global top tier
HiDream.ai released HiDream-O1-Video-1.0 on September 15, describing it as the first natively omni-modal video generation model, with text, image and video inputs producing 1080p clips of five to twenty seconds. It ranked fourth on the Artificial Analysis Image to Video Leaderboard (With Audio) and eighth on Arena.ai's image-to-video blind evaluation, while the company announced a C+ round backed by Newmicro Capital, Jiaozi Capital and ICBC Capital.
HiDream.ai on September 15 formally released HiDream-O1-Video-1.0, which it describes as the first natively omni-modal video generation model. According to the company's official announcement carried by QbitAI, the model, abbreviated HD-V1, is built on HiDream's in-house omni-modal architecture and accepts text, image and video inputs, generating 1080p clips of five to twenty seconds in a single pass.
The launch came with two third-party benchmark placements. HD-V1 ranked fourth globally on the Artificial Analysis Image to Video Leaderboard (With Audio), and eighth on Arena.ai's Image-to-Video blind evaluation. Artificial Analysis is widely cited for running standardized, reproducible comparisons of leading video models, while Arena.ai's ranking relies on blind testing, making both commonly used reference points for judging image-to-video capability.
Architecturally, HD-V1 places a multimodal intent-understanding module ahead of generation. That module converts loose, colloquial user prompts into a structured plan covering shot duration, location, visual content, first-frame constraints, characters present, action and expression, framing and lens choice, plus dialogue and background audio design. Only after that planning step does the model move into joint generation, an approach the company says reduces missed intent, jarring shot changes and character drift.
Physics is the second pillar the company emphasizes. How objects fall under gravity, how collisions respond, how inertia carries, how light decays and how space remains continuous are described as underlying logic written into frame generation. In side-by-side demos released with the announcement, comparison models produced physically implausible results such as a taut line shortening as two hands pulled inward or a needle appearing from nowhere, while HD-V1 kept the deformation, tension and direction of the line consistent with the force being applied.
On audio, HD-V1 takes a natively omni-modal route and models text, video and audio signals jointly so the three constrain and align with one another inside a single generation process, rather than generating frames first and dubbing sound afterwards. The company says its post-training stage uses Diffusion Reinforcement Learning together with a multimodal reward model aligned to human perception, scoring content from picture to sound as a whole. Clip length is also decided by the content: the model plans duration from the logic of the event and the rhythm of the narrative instead of a hard-coded number of seconds in the prompt.
Seen against HiDream's own lineup, the video model closes a loop. Since unveiling the HiDream-O1 native omni-modal world-model architecture in April, the company has shipped a family of models on the same architecture: the commercial version of HiDream-O1-Image 1.5 placed first in China and second globally on Artificial Analysis' text-to-image leaderboard, the interactive world model HiDream-O1-World topped the overall WBench ranking, and the embodied world model HiDream-O1-Embodied led sub-leaderboards on disturbance adaptation and spatial reasoning in RoboColiseum. With the video model out, the company says the industry's first native omni-modal world-model loop is complete.
The business side moved on the same day. HiDream announced a C+ funding round alongside the launch, with Newmicro Capital, Jiaozi Capital and ICBC Capital participating. Newmicro Capital, backed by Shanghai Xinshwei Technology Group, Shanghai Science and Technology Investment and its management team, manages close to 10 billion yuan, while Jiaozi Capital is a wholly state-owned equity investment platform under Chengdu Jiaozi Financial Holdings with more than 170 billion yuan under management. Both framed the investment around HiDream's world-model matrix and their own industrial resources and computing infrastructure.
For HiDream, the release shifts the company from isolated capabilities toward a single foundation. Founder Mei Tao has described the strategy as one native omni-modal base with four models evolving from the same source, while CTO Yao Ting called HD-V1 a key part of the world-model matrix and said the twin leaderboard results are direct validation of that technical path. What to watch next is how quickly the video model opens to users, whether its rankings hold, and how fast state-backed capital translates into computing capacity and industry deployments.
Why it matters
The release completes the video layer of HiDream's world-model lineup and shows a Chinese lab competing on the same leaderboards as global front-runners in image-to-video. State-backed capital joining the round could speed up its computing build-out and industry deployments.
Nearby Updates
All09/15, 18:57
Perplexity brings its local Portable Computer agent to Windows PCs with RTX GPUs
Perplexity made its on-device Portable Computer agent available in the Windows app on September 14, running entirely on NVIDIA GeForce RTX or RTX PRO GPUs with at least 24GB of VRAM and open to Pro and Max subscribers. Unlike the cloud version, the model, agent harness, orchestrator and scheduler all run on the machine, so sensitive data stays on the device and locally completed work does not consume Perplexity Computer credits.
09/15, 17:58
MoleculeMind pushes QuantaMind to 100,000-atom reaction simulations on a single GPU
MoleculeMind says its reactive machine-learning force field QuantaMind can now simulate reactions in 100,000-atom systems for hundreds of nanoseconds at near quantum-chemistry accuracy, at about 0.25 seconds per step on a single GPU. The underlying Science Advances paper ran a continuous 6-nanosecond simulation of a 17,792-atom PETase system, covering proton transfer, bond breaking and formation and the full catalytic cycle, with agreement above 0.99 against quantum-mechanical checks.
09/15, 17:50
MediaTek Launches 2nm Dimensity 9600 Pro Flagship, Betting on an AI-Native Architecture for Agents
MediaTek has unveiled the Dimensity 9600 Pro, a 2nm flagship mobile chip marketed around an AI-native architecture and aimed explicitly at agentic AI. It is a clear signal that on-device AI is shifting from “can it run a model” to “can it run agents well, continuously.”
09/15, 17:43
Zidong Taichu open-sources ZDTaichu5.0-9B, pitching spatial embodiment under 10B parameters
The Zidong Taichu series has open-sourced ZDTaichu5.0-9B, which the release describes as the strongest general multimodal model under 10 billion parameters for spatial embodied ability. Keeping the model at the 9B level points at a clear goal: getting multimodal understanding onto robots and other physical devices rather than chasing general chat leaderboards.