Frontier capability
How fast is frontier cyber capability rising?
Narrow tasks: apprentice-level success rate
Published points only.
Cyber ranges: full completions out of 10
Published points only.
Global peer evidence and every published point
Late 2023 · Best tested frontier model
Apprentice-level cyber task success: 9%
Caveat: Lab task success rates on narrow tasks, not real-world attack rates. AISI notes the trend is not guaranteed to continue.
Mar 2025 · Frontier agents
HCAST: 70 to 80% success on tasks taking skilled humans under one hour
Caveat: Midpoint shown only to register the published 70 to 80% range. HCAST combines machine learning, cyber, software and reasoning tasks, so this is not a cyber-only score and is not comparable with AISI points.
Jun 2025 · Top-performing agent combinations
CyberGym: about 20% success on vulnerability reproduction tasks
Caveat: Benchmark result for reproducing known vulnerabilities from source and descriptions. It is not a real-world exploitation rate and is not comparable with AISI task success.
2025 · First model to do so
First expert-level (10+ years experience) cyber task completions
Caveat: A milestone, not a rate. No success percentage is published for this point, so none is shown. Plotted at mid-2025 because AISI reports the year only.
Sep 2025 · Gemini frontier models
Frontier Safety Framework defines tracked cyber capability levels
Caveat: A risk framework and threshold definition, not a comparable published cyber score. No numeric point is plotted.
Late 2025 · Best tested frontier model
Apprentice-level cyber task success: 50%
Caveat: Lab task success rates on narrow tasks, not real-world attack rates. AISI notes the trend is not guaranteed to continue.
Apr 2026 · Claude Mythos Preview
Mythos Preview: first model to complete The Last Ones, 3 of 10 attempts
Caveat: Cyber ranges lack active defenders. A 32-step simulated corporate network estimated at around 20 expert hours.
Apr 2026 · GPT-5.5
GPT-5.5 completed The Last Ones in 2 of 10 attempts
Caveat: Cyber ranges lack active defenders. A 32-step simulated corporate network estimated at around 20 expert hours.
May 2026 · Claude Mythos Preview (newer checkpoint)
Newer Mythos Preview checkpoint: The Last Ones 6 of 10
Caveat: Small, undefended enterprise networks, where initial access has already been gained.
May 2026 · Claude Mythos Preview (newer checkpoint)
First ever completion of the Cooling Tower ICS range: 3 of 10
Caveat: Simulated industrial control range without active defenders. Not evidence of real-world OT compromise.
Jul 2026 · Claude Opus 5
Lab system card reported a broader cyber evaluation suite
Caveat: The system card reports internal and external cyber evaluations, but this monitor does not plot an unverified or incomparable numeric score.
2026 system card · GPT-5.3-Codex
Lab system card treated the model as High cyber capability under its framework
Caveat: Categorical, precautionary threshold treatment. The source says it did not have definitive evidence that the model reached the threshold, so no numeric value is shown.
Sep 2026 · Leading models at snapshot
Artificial Analysis Cyber Index leaders scored 56
Caveat: Live leaderboard snapshot. The 0 to 100 composite covers defensive vulnerability finding and patching only. It is not a task success percentage, attack rate or metric comparable with AISI.
Last point: 28 Sept 2026.
Glossary: what the evaluations measure
Narrow tasks
Short capture-the-flag style challenges that test one skill, such as reverse engineering, web exploitation or cryptography. AISI groups them by difficulty, from technical non-expert to apprentice and expert level. Good for tracking skills over time; they say little about running a whole intrusion.
Cyber ranges
Simulated networks with many hosts and a chain of steps to reach a goal. "The Last Ones" is a 32-step corporate network; "Cooling Tower" is an industrial control system range. Closer to a real intrusion, but there are no active defenders and initial access is typically already given.
What this does and does not show
- Capability on published lab evaluations is rising quickly.
- More than one developer's model has now completed a multi-step range.
- The first ICS range completion was published in May 2026.
- Peer benchmarks broaden coverage without being combined into a synthetic global score.
- It does not show real-world attack rates or success against defended networks.
- It does not show that OT is already under AI-driven attack.
- It is not a forecast. AISI notes trends are not guaranteed to continue.