HeadlinesBriefing favicon HeadlinesBriefing.com

GLM-5.3: Frontier Coding with Emergent Cyber Capabilities

Hacker News •
×

Z.ai released GLM-5.3, achieving a 50% improvement over GLM-5.2 on the in-house Z.ai Code Bench and claiming open-source SOTA on Terminal Bench 3.0 and Agents' Last Exam. Every gain comes from scaling post-training on the same base model, using expanded long-horizon environments and the SAO RL stack.

On coding benchmarks, GLM-5.3 jumps from 4.6 to 28.3 on Terminal Bench 3.0 and from 46.2 to 66.9 on Deep SWE v1.1. The Z.ai Code Bench shows 34.5% task completion at Max effort with 75K output tokens, surpassing GLM-5.2's 23.4% at 96K tokens and exceeding Claude Opus 4.8 at High effort.

Emergent cyber capability developed faster than expected. GLM-5.3 scores 84.5% on Cyber Gym (vulnerability discovery), leads on Exploit Bench at 54.4% (more than doubling GLM-5.2's 24.4%), and completes 105/130 tasks on Exploit Gym within 2/6 hours versus 29/39 for GLM-5.2. Gains are largest further up the exploitation chain.

Weights will be open-sourced in two weeks after safety evaluation. Training used synthesized environments from real engineering workflows, with verifier agents ensuring executable, verifiable tasks and binary rewards for RL.