UK AI Security Institute finds open-weight models now trail closed frontier on cyber by months
The UK AI Security Institute found open-weight models like GLM-5.2 now trail closed frontier models on cyber tasks by four to seven months, down from six to 10.
The UK AI Security Institute said on July 17, 2026 that leading open-weight models now trail closed frontier models on cyber-capability tasks by four to seven months, down from a six-to-10-month gap a year earlier.
The finding, published by the UK government’s model-testing agency, quantifies how fast freely downloadable models are catching the best closed systems on offensive-cyber tasks — the capability regulators worry about most, because open weights cannot be recalled once released. AISI tested two open-weight leaders: GLM-5.2, released in June 2026, and DeepSeek V4-Pro.
On a 70-task narrow-cyber benchmark, GLM-5.2 matched Anthropic’s Claude Opus 4.6 (a four-month lag) and, on AISI’s ‘The Last Ones’ cyber range, reached as far as the older Claude Opus 4.5 (a seven-month lag). DeepSeek V4-Pro compared to Opus 4.5 at a five-month lag while costing about $0.28 per task against $12.50 for that closed model; GLM-5.2 ran at $6.12 per task versus $15.17 for Opus 4.6, according to AISI’s analysis.
The results measure narrow cyber benchmarks, not real-world attacks. A closing capability gap does not mean open-weight models can run end-to-end intrusions unaided; benchmark performance and operational offensive capability are not the same thing.
Still, the direction is what matters for policy. If the gap keeps narrowing, the window in which closed-model safeguards can meaningfully gate dangerous capabilities shrinks with it.
More news

AWS releases six open-source Hugging Face deployment skills for SageMaker

Google Research releases MilleMiglia logistics benchmark generator

AWS launches AgentCore Runtime V2 with elastic memory and snapshot starts
