Open-Weight Models Are Now 4 to 7 Months Behind the Cyber Frontier. The UK Just Put a Number on It.
TL;DR
The UK's AI Security Institute (AISI) ran the leading open-weight models through its offensive-cyber evaluations and answered a question every red team and every defender has been circling: how much of a head start does keeping your best model closed actually buy you? The answer is 4 to 7 months. GLM-5.2 and DeepSeek V4-Pro now perform on cyber tasks like the closed frontier did 4 to 7 months earlier, down from a 6 to 10 month gap through most of 2025. The lead is real, but it is measured in a couple of model cycles, not years, and it is still shrinking.
The Head Start Has a Number Now
Until this week the "open models are catching up on security" conversation ran on vibes. AISI, the UK government body that stress-tests frontier systems, just replaced the vibes with a measurement. Its verdict, published July 17: the most cyber-capable open-weight models trail the closed frontier by roughly 4 to 7 months, versus the 6 to 10 months AISI logged internally through most of 2025.
The specifics are sharper than the headline range. GLM-5.2, released in June 2026, lands next to Anthropic's Opus 4.6 (February 2026) on AISI's short cyber tasks and next to Opus 4.5 (November 2025) on the harder end-to-end scenarios. Put differently, an open model you can download today roughly matches what the closed frontier, including OpenAI's GPT-5.3-Codex, could do about four months ago. DeepSeek V4-Pro sits a touch further back, comparable to Opus 4.5 from five months prior.
Two Ways to Measure a Hacker
AISI does not grade cyber skill with one number, because "can it hack" is really two questions. The first is whether a model knows the individual moves. The second is whether it can string them together into a real break-in without a human holding its hand. So the institute runs two very different tests.
Narrow cyber tasks probe specific skills across four difficulty levels: vulnerability research, exploitation, reverse engineering, web exploitation, and cryptography. AISI ran a 70-task subset, five attempts per task, with a 2.5-million-token budget. Think of this as the written exam: discrete problems with a right answer.
Cyber ranges are the practical. They drop a model into a simulated corporate network of hosts, services, and planted vulnerabilities, and ask it to run the whole intrusion autonomously. One range, ominously named "The Last Ones," is a 32-step attack across a corporate network that AISI estimates would take a human expert about 20 hours. This is less a quiz and more a flight simulator for breaking into a company: the model has to fly the entire mission, not just answer trivia about aerodynamics. Models that ace the narrow tasks routinely stall on the ranges, which is exactly why AISI reports both.
An Independent Read Lands in the Same Place
AISI is not the only lab that put DeepSeek V4-Pro on the bench. An independent evaluation from Neo Research pegs the model at three to six months behind the Western frontier on cyber, right in line with AISI's read. On Cybench, the Stanford capture-the-flag benchmark that the US and UK safety institutes both fold into their own frameworks, the numbers are close enough to make the point.
Neo Research notes DeepSeek is weaker on the fiddly, sustained-reasoning stuff, smart-contract exploitation on EVMbench being a soft spot. That matches AISI's split: parity on the isolated skills, a lag on the long autonomous chains. The frontier's remaining moat is stamina, not knowledge.
Why 4 Months Is the Whole Point
AISI is refreshingly blunt about why it measures this at all. The gap is not a scoreboard, it is a clock. In the institute's words, the distance between open and closed cyber capability "provides a preparation time: a window for cyber defenders with access to the most capable closed systems to take action before today's frontier cyber capabilities might become available without the same safeguards."
Translation: whatever an open-weight model can do on offense today, a defender with frontier access has already seen for months. When the gap was 6 to 10 months, that was a comfortable buffer to patch, detect, and harden. At 4 months it is a sprint, and closing. The open frontier is essentially speedrunning the closed one, which is great news if you build with open weights and a slightly worse night's sleep if you defend networks for a living.
This tracks AISI's broader Frontier AI Trends Report, which puts the general open-to-closed capability gap at roughly 4 to 8 months (an Artificial Analysis-style intelligence index shows about 4 months, METR's time-horizon benchmark about 8 as an upper bound). Cyber, in other words, is not a special case racing ahead. It is moving in lockstep with raw capability, which is the more sobering read.
What This Means If You Build With Open Weights
- Capability is no longer the differentiator you assumed. If your threat model treated open-weight models as toys that lag the frontier by a year, requantify. On isolated cyber skills, GLM-5.2 and DeepSeek V4-Pro are already frontier-adjacent.
- The moat is autonomy, for now. Both open models still trip on long, multi-step intrusions. If you rely on that gap, watch it, because it is the thing shrinking fastest.
- Defenders should treat the buffer as expiring. Four months of preview is not a plan. Assume open-weight parity on offense sooner than the last cycle taught you.
- The next data point is imminent. AISI says Kimi K3 weights land at the end of July, and it plans to test them. A 2.8-trillion-parameter open model could move this line again within weeks.
Key Takeaways
- AISI measured leading open-weight models on offensive cyber and found them 4 to 7 months behind the closed frontier, down from 6 to 10 months in 2025.
- GLM-5.2 matches Opus 4.6 (Feb 2026) on narrow tasks and Opus 4.5 (Nov 2025) on end-to-end ranges; DeepSeek V4-Pro is comparable to Opus 4.5.
- AISI grades two ways: narrow skill tasks (the exam) and cyber ranges (the flight simulator), and open models lag most on the long autonomous chains.
- Independent Neo Research puts DeepSeek V4-Pro three to six months back, and on Cybench it matched GPT-5.2 and beat Opus 4.5.
- The gap functions as defender "preparation time," and at 4 months that buffer is thin and still closing.
- Kimi K3's open weights, due end of July, are next up for AISI testing and could shift the line again.
Sources: AISI: How Far Behind the Frontier are Leading Open Weight Models on Cyber?, AISI Frontier AI Trends Report, Neo Research: Evaluating DeepSeek V4 Pro for Frontier Risks, Cybench, NIST CAISI: Evaluation of DeepSeek V4 Pro