Frontier capability

How fast is frontier cyber capability rising?

AISI provides the comparable UK spine. Global peer evaluations appear below in their own series because unlike benchmarks cannot be merged. All are lab results, not real-world attack rates.

Narrow tasks: apprentice-level success rate

Published points only.

Cyber ranges: full completions out of 10

Published points only.

Global peer evidence and every published point

  1. Late 2023 · Best tested frontier model

    Apprentice-level cyber task success: 9%

    Caveat: Lab task success rates on narrow tasks, not real-world attack rates. AISI notes the trend is not guaranteed to continue.

  2. Mar 2025 · Frontier agents

    HCAST: 70 to 80% success on tasks taking skilled humans under one hour

    Caveat: Midpoint shown only to register the published 70 to 80% range. HCAST combines machine learning, cyber, software and reasoning tasks, so this is not a cyber-only score and is not comparable with AISI points.

  3. Jun 2025 · Top-performing agent combinations

    CyberGym: about 20% success on vulnerability reproduction tasks

    Caveat: Benchmark result for reproducing known vulnerabilities from source and descriptions. It is not a real-world exploitation rate and is not comparable with AISI task success.

  4. 2025 · First model to do so

    First expert-level (10+ years experience) cyber task completions

    Caveat: A milestone, not a rate. No success percentage is published for this point, so none is shown. Plotted at mid-2025 because AISI reports the year only.

  5. Sep 2025 · Gemini frontier models

    Frontier Safety Framework defines tracked cyber capability levels

    Caveat: A risk framework and threshold definition, not a comparable published cyber score. No numeric point is plotted.

  6. Late 2025 · Best tested frontier model

    Apprentice-level cyber task success: 50%

    Caveat: Lab task success rates on narrow tasks, not real-world attack rates. AISI notes the trend is not guaranteed to continue.

  7. Apr 2026 · Claude Mythos Preview

    Mythos Preview: first model to complete The Last Ones, 3 of 10 attempts

    Caveat: Cyber ranges lack active defenders. A 32-step simulated corporate network estimated at around 20 expert hours.

  8. Apr 2026 · GPT-5.5

    GPT-5.5 completed The Last Ones in 2 of 10 attempts

    Caveat: Cyber ranges lack active defenders. A 32-step simulated corporate network estimated at around 20 expert hours.

  9. May 2026 · Claude Mythos Preview (newer checkpoint)

    Newer Mythos Preview checkpoint: The Last Ones 6 of 10

    Caveat: Small, undefended enterprise networks, where initial access has already been gained.

  10. May 2026 · Claude Mythos Preview (newer checkpoint)

    First ever completion of the Cooling Tower ICS range: 3 of 10

    Caveat: Simulated industrial control range without active defenders. Not evidence of real-world OT compromise.

  11. Jul 2026 · Claude Opus 5

    Lab system card reported a broader cyber evaluation suite

    Caveat: The system card reports internal and external cyber evaluations, but this monitor does not plot an unverified or incomparable numeric score.

  12. 2026 system card · GPT-5.3-Codex

    Lab system card treated the model as High cyber capability under its framework

    Caveat: Categorical, precautionary threshold treatment. The source says it did not have definitive evidence that the model reached the threshold, so no numeric value is shown.

  13. Sep 2026 · Leading models at snapshot

    Artificial Analysis Cyber Index leaders scored 56

    Caveat: Live leaderboard snapshot. The 0 to 100 composite covers defensive vulnerability finding and patching only. It is not a task success percentage, attack rate or metric comparable with AISI.

Last point: 28 Sept 2026.

Glossary: what the evaluations measure

Narrow tasks

Short capture-the-flag style challenges that test one skill, such as reverse engineering, web exploitation or cryptography. AISI groups them by difficulty, from technical non-expert to apprentice and expert level. Good for tracking skills over time; they say little about running a whole intrusion.

Cyber ranges

Simulated networks with many hosts and a chain of steps to reach a goal. "The Last Ones" is a 32-step corporate network; "Cooling Tower" is an industrial control system range. Closer to a real intrusion, but there are no active defenders and initial access is typically already given.

What this does and does not show

  • Capability on published lab evaluations is rising quickly.
  • More than one developer's model has now completed a multi-step range.
  • The first ICS range completion was published in May 2026.
  • Peer benchmarks broaden coverage without being combined into a synthetic global score.
  • It does not show real-world attack rates or success against defended networks.
  • It does not show that OT is already under AI-driven attack.
  • It is not a forecast. AISI notes trends are not guaranteed to continue.