HeadlinesBriefing favicon HeadlinesBriefing.com

GLM-5.3-Flash: Frontier AI at Flash Cost

Hacker News •
×

2026-08-26 · Research GLM-5.3-Flash: Frontier Intelligence, Flash Cost* Call it at Z.ai* Z.ai Coding Plan* Code with ZCode* Hugging Face We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.

GLM-5.3-Flash incorporates several architectural improvements over GLM-5. For the first time, we introduce a hybrid architecture combining sparse and linear attention, sharply reducing long-context serving costs while preserving precise long-context capabilities. It also adopts Manifold-Constrained Hyper-Connections (m HC) to further improve scaling efficiency. Combined with our latest 30T-token multimodal pre-training corpus, these changes let GLM-5.3-Flash produce more intelligence with less compute.

Before release, we tested GLM-5.3-Flash anonymously as `ox-alpha` on Open Code and Open Router to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips. GLM-5.3-Flash pushes the Pareto frontier of the Artificial Analysis Intelligence Index v4.1.1, scoring 57 at just $0.045 per task (discounted) — a level of intelligence previously only available at roughly 10× the cost.

Across six coding and agentic benchmarks, GLM-5.3-Flash consistently outperforms GLM-5.2, often by a wide margin — 63.4 vs. 46.2 on Deep SWE v1.1 and 48.8 vs. 26.2 on Automation Bench — while approaching Claude Opus 4.8 overall. Compared with the GLM-4.5 series, GLM-5.3-Flash is specifically designed for ultra-low-cost inference, nearly halving both the activated parameter count and the number of layers.