Next upHack for Humanity: San Francisco (powered by Google Gemini)
News

Z.ai ships GLM-5.3, topping the CyberGym security benchmark without a retrain

Z.ai released GLM-5.3 on Aug. 14, scoring 84.5 on the CyberGym vulnerability-discovery benchmark using only post-training scaling of its GLM-5.2 base model.

D
Aug 14, 2026 · 1 min read

Z.ai (Zhipu AI) released GLM-5.3 on Aug. 14, a large language model that scored 84.5 on the CyberGym vulnerability-discovery benchmark, ahead of Claude Mythos 5 and GPT-5.6 Sol, on the company’s own numbers.

What stands out is how Z.ai built it. The company said GLM-5.3 reuses the GLM-5.2 base model with no retraining, crediting the entire capability jump to scaled-up post-training. On the Terminal-Bench 3.0 agentic benchmark, the model rose to 28.3 from GLM-5.2’s 4.6; on DeepSWE v1.1 it moved to 66.9 from 46.2.

The security results are the sharper story. Z.ai said the model, trained on vulnerability-discovery data, began reasoning across multiple stages of exploitation and assembling complete exploit chains rather than isolated bug-finds. Since GLM-5.2, it has flagged 2,436 vulnerabilities across 269 open-source projects, 1,097 of them rated critical or high severity, according to the company’s announcement.

Those figures are Z.ai’s own and have not been independently verified; the CyberGym margins over Claude Mythos 5 (83.8) and GPT-5.6 Sol (83.6) are narrow. Benchmark scores also translate unevenly to real-world exploitation.

The dual-use tension is explicit. GLM-5.3 is available now through the GLM Coding Plan, with API access coming soon. Z.ai said it is holding back the open weights for roughly two weeks for safety evaluation and hardening before release.

Whether that review changes what ships — or whether the weights arrive unmodified — is the signal worth watching for anyone tracking how far open models push offensive security capability.

More news