Realtime AI News
MiniMax Releases M3 Multimodal Model Series with Base and Quantized Versions
MiniMax has published the M3 multimodal model series on Hugging Face, including a base model and the MXFP8 quantized version. The series supports image-text-to-text tasks, uses MoE architecture, and offers agent and coding capabilities.

MiniMax has officially released the M3 multimodal model series on Hugging Face, featuring both the base M3 model and the quantized MiniMax-M3-MXFP8. The models adopt a Mixture of Experts (MoE) architecture, supporting image-text-to-text multimodal tasks and offering agent and coding capabilities.
According to the Hugging Face page, the M3 series quickly gained traction, with over 572,000 downloads and 43 likes. This reflects strong community demand for open-source multimodal models.

The key highlight of the MiniMax M3 series is its multimodal fusion capability, processing both image and text inputs to generate text outputs. This makes it suitable for visual question answering, image-text generation, intelligent customer service, and more.
The quantized MXFP8 version uses 8-bit floating-point quantization to reduce deployment cost, enabling large model inference on consumer-grade GPUs and lowering the barrier for developers.
MiniMax has been committed to large model R&D, and the M3 series represents a significant step in multimodal direction. Compared to similar models, M3 strikes a good balance between parameter count and performance.
Looking ahead, MiniMax may release larger models or optimize for specific scenarios. Developers can expect richer toolchains and community support.
The release also underscores the growing activity of Chinese AI companies in the open-source space, offering more choices to global developers.
Why it matters
MiniMax's M3 multimodal series is a notable open-source contribution, advancing multimodal model accessibility and cost-effective deployment.
Nearby Updates
All07/01, 12:36
MiniMax Officially Releases M3 Multimodal MoE Model on Hugging Face with Image-Text-to-Text Pipeline
MiniMax has officially released its M3 model on Hugging Face, a multimodal mixture-of-experts (MoE) model supporting an image-text-to-text pipeline. The release has already garnered over 192,000 downloads and 1,271 likes on the platform.
07/01, 13:01
First AI Agent Payment Completed in France, Marking a FinTech Milestone
France has completed the first payment transaction initiated and executed autonomously by an AI agent, as reported by FinTech Futures. The event marks a critical step for AI agents moving from information processing to real-world financial operations.
07/01, 12:03
GSMA Intelligence Releases Agentic Core White Paper, Defining New Paradigm for Intelligent Core Network Evolution
GSMA Intelligence has released the Agentic Core white paper, defining a new paradigm for the evolution of intelligent core networks. The document aims to guide the telecom industry towards AI-driven network architectures.
07/01, 11:46
Om AI Lianhui Releases VLX: World's First Edge Streaming Multimodal Model for the Physical World
Om AI Lianhui has unveiled VLX, claiming it as the world's first edge streaming multimodal model designed for the physical world. The model enables real-time multimodal processing on edge devices, reducing latency and enhancing privacy.