Guozhen AIGlobal AI field notes and model intelligence

Realtime AI News

Zhipu unveils GLM-5.3 Flash: the mystery 'Ox Alpha' model is its first natively multimodal open-source GLM, served on domestic chips

Zhipu has confirmed that the viral mystery model 'Ox Alpha' is GLM-5.3 Flash, its first natively multimodal model in the GLM 5 series, released as open source and served entirely on domestic Chinese chips. With 320B total parameters and 18B active, it beats the larger GLM-5.2, scores 57 on the AA benchmark tied with Claude Opus 4.8, and is priced at 1/40 of Opus 4.8.

Published
「牛来」真身曝光:智谱开源GLM首个原生多模态模型GLM-5.3 Flash,跑在国产卡上
Image source: qbitai.com

The mystery model "Ox Alpha" that has been sweeping overseas developer circles finally has a name: it is GLM-5.3 Flash, Zhipu AI's newly released and immediately open-sourced model, and the first natively multimodal model in the GLM 5 series.

Before Zhipu claimed it, the anonymous Ox Alpha hit the top of OpenRouter on its first day and set a record for single-day token usage; on OpenCode, its arrival ended DeepSeek's 56-day run at the top of the leaderboard.

With 320B total parameters, GLM-5.3 Flash outperforms the larger GLM-5.2 (753B). It scores 57 on the latest AA benchmark, tied with Claude Opus 4.8, while its pricing comes in at 1/10 of GLM-5.3, with a limited-time discount down to 1/20 of GLM-5.3 and 1/40 of Opus 4.8 — also lower than DS-V4-Flash.

Under the hood, GLM-5.3 Flash activates only 18B parameters, compresses layers from 92 (GLM-4.5 era) to 45, and uses a hybrid linear-attention plus sparse-attention architecture. With an IndexPool that reduces the indexer's cache vectors from four to one, attention computation drops 3.01x and KV cache shrinks 4.44x compared with GLM-5.3.

Zhipu also built a data-synthesis pipeline for Visual Coding, letting the model look at final pages, interactions, and 3D scenes during a task and revise based on visual feedback. In an official demo, GLM-5.3 Flash ran for 12 hours on its own to build a 3D Blender scene.

Even more striking is the compute story. The 62T tokens of real global traffic during anonymous testing were all served on domestic Chinese accelerator chips. Zhipu split multimodal encoding, prompt prefill, and token-by-token decoding into an independently scalable Encode-Prefill-Decode architecture, layering on optimizations like Layer Split and mixed cache quantization to triple end-to-end serving performance and bring per-token cost in line with mainstream Nvidia GPUs.

QbitAI's hands-on tests found GLM-5.3 Flash can add precise subtitles to multi-speaker videos, turn an entire movie into a narrated explainer, and turn a single UI design mockup into an interactive shopping app — demonstrating a write-look-edit loop.

GLM-5.3 Flash is now open source worldwide, integrated with ZCode and open APIs, with weights on Hugging Face and self-deployment options for enterprises. Model, domestic compute, and the open-source ecosystem are finally connected — whether it can keep delivering frontier-grade results at low cost is the question to watch.

Why it matters

By combining native multimodality, open weights, and domestic chip serving at a fraction of frontier prices, GLM-5.3 Flash could accelerate both price deflation and the open-source shakeup in frontier AI.

ZhipuGLMOpen Source
Back to realtime news

Nearby Updates

All