Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

MiniMax launches omnimodal model H3 with 15-second 2K video at under a third of mainstream pricing

MiniMax has officially released MiniMax H3, a universal omnimodal generation model that unifies text, image, video and audio understanding and generation, producing up to 15-second 2K audio-video with native dual-channel sound. The company says per-second generation cost at 2K is under a third of mainstream models, and it plans to open the model's weights within days, subject to regulations.

Published
MiniMax发布全模态生成模型H3:支持15秒2K音视频,价格不到主流三分之一
Image source: minimax.io

AI company MiniMax has officially released MiniMax H3, a universal omnimodal generation model. The company says H3 is the first of its kind to unify understanding and generation across text, image, video and audio, breaking down the silos between tasks and modalities in video generation.

H3 outputs audio-video with native dual-channel sound, supporting up to 15-second clips at 2K resolution. In early invite-only testing, the model showed commercial-grade stability in instruction following, brand-information rendering and V2V action transfer, positioning it for advertising, e-commerce, gaming and UI design.

Users describe complex relationships in natural language — for example, referencing the Hitchcock-style camera movement of a video and asking the model to make a person in an image sing a certain audio track — and the model handles the full pipeline of understanding and generation automatically.

On the technical side, H3 introduces Contextual Omni Representation, which turns natural language into a universal bridge across modalities. Combined with the in-house H3-VAE and in-context regeneration techniques, per-second generation cost at 2K comes to less than a third of mainstream models, and under half at 768P.

MiniMax also said it will open H3's model weights within the coming days, subject to relevant laws and regulations — the first time a Chinese video-generation model has opened its weights to the community, a move aimed at boosting the open-source ecosystem and accelerating adaptation of domestic chips, for which H3 was designed with compatibility in mind.

MiniMax acknowledged that H3 still has room to improve in image fidelity and model scale, and said it will push toward fusing the H-series and M-series capabilities while continuing to scale the model's ceiling.

Industry analysts expect the release and open-sourcing of H3 to challenge the closed-source dominance in video generation and accelerate the democratization of multimodal AI.

Why it matters

H3's aggressive pricing and promised open weights could reshape China's video-generation market and speed up domestic chip adaptation. Watch for the weight release in the coming days and how the H-series merges with MiniMax's M-series.

MiniMaxH3Video Generation
Back to realtime news

Nearby Updates

All