Cyber Defense Is Losing
While at DARPA I once debated that Cybersecurity would at some point be solved. I argued that if we just built and implemented software correctly, DEFCON would be no more exciting than whatever the people who make bolts and hardware do for a convention. Boring, commoditized and expected.
Well, I’ll concede that cyber isn’t getting solved anytime soon. While AI may eventually present the ability to develop and implement secure software, right now the attacker is getting nearly all the benefits. This is because AI generated code is buggy, our world is increasing in scale and complexity, and now offense is easy while defending remains hard.
For a long time, there has been a fragile balance between offense and defense — a perfect environment to build a cyber company. Everything has been getting bigger and more digital at once: more data, more devices, more software, more dependencies. Complexity and scale make defense difficult, but new tech always rose to this challenge. This all worked as long as offense remained the domain of the skilled hacker.
Recent AI advances turn two known properties of the offense/defense game into a big problem at scale, tipping this balance. The life blood of cyber is vulnerability discovery, we learn how software can break. This information allows attackers to write exploits to take over systems, while defenders implement patches to fix the software.
First, while both are fragile, exploits are certain while patching is an exercise in hope. Once an attacker has working exploit code, it tends to work the same way every time against every still-vulnerable target. The attacker has a master key. The defender, by contrast, has to run a messy social and technical process: identify the affected systems, test the fix, schedule the outage, avoid breaking production, coordinate with vendors, and hope nothing critical was missed. You never really know how certain a patch is. The exploit (and the access it gives you) is truth. The patch is a potential improvement in security.
Second, offense needs one opening while defense needs universal coverage. A large enterprise is not defending a single castle wall. It is defending an every changing city with millions of doors: endpoints, cloud workloads, identities, APIs, SaaS integrations, vendor pathways, forgotten development environments, and software libraries nested inside software libraries. The attacker only needs one of those doors to stay open long enough to matter.
Those two facts explain more about the current cyber moment than most product categories, frameworks, or buzzwords. And both of them are radically accelerated by the long trend that mythos is putting right in front of us.

The Patch Gap
Think of a flaw in a widely used lock. The moment someone publishes a working bypass, every copy of that lock in the world shares the same weakness. The offensive advantage arrives instantly and spreads perfectly. The defensive fix does not.
That fix has to diffuse through institutions. Some teams patch in hours. Some patch in weeks. Some wait for a vendor. Some do not know they are exposed. Some know, but cannot patch without risking a more immediate outage. Some are carrying vulnerable code inside a container inside a product they do not control. The result is the same pattern every time: attacker capability jumps to near total coverage immediately, while defender coverage crawls upward over months or years and never quite reaches 100 percent. Additionally, the patch will prevent the known mechanism of the exploit, but do other vulnerabilities exist? Maybe the new patch created them.
This is the patch gap. It is where a large fraction of real-world compromise lives.
Real example: Log4Shell made the pattern visible to everyone. The vulnerability was obvious, the exploitation path was clear, and the world still spent months chasing it. Years later, defenders were still seeing exploit attempts because the tail never fully disappeared. The lesson was not just that organizations should patch faster. It was that patching is constrained by logistics, testing, ownership, incentives, and plain institutional friction. Even when everyone agrees on what should happen, it does not happen uniformly.

AI makes this worse before it makes it better. It is already helping discover vulnerabilities faster, write proof-of-concept exploits faster, and scale reconnaissance faster. That compresses the time between flaw discovery and usable offensive tradecraft. Unless defensive remediation becomes equally automated, the patch gap widens.
One vs. All
The second asymmetry is simpler and harsher. Attackers need one success. Defenders need comprehensive performance across an enormous attack surface.
There is a reason experienced defenders sound pessimistic even when they are competent. They have internalized the arithmetic. If an organization has \(n\) things that can fail, and each of them is secure with probability \(p\) on a given day, then the chance that everything is secure is \(p^n\). Once \(n\) is large enough, even very high values of \(p\) stop feeling reassuring.
That is not a slogan. It is a scaling law.
Suppose your controls are extremely good. Suppose each asset is secure 99.999 percent of the time. That sounds excellent until you apply it across millions of assets, identities, services, and dependencies. The question stops being whether a failure exists and becomes where it exists, how exposed it is, and whether the attacker finds it before you do.
This is why “we are pretty good on average” is not a comforting statement in cyber defense. Averages do not defend networks. Coverage does.
It also explains why the workforce problem is not simply a hiring problem. The surface being defended has grown much faster than the workforce available to defend it. Cloud expansion, SaaS sprawl, API dependency chains, machine identities, and AI-generated code have all increased the number of doors faster than institutions can add trained defenders. Even if every hiring pipeline worked better, the denominator is still outrunning the numerator.
Why AI Changes the Balance
AI does not create these asymmetries, but it does intensify them. The core question is “what can AI do and what permissions do we give it?” It can now exploit the software it’s given. It can (and should not) have access to all computers and the authority to quickly patch them.
On offense, AI lowers the cost of attacks that used to be too expensive to bother with. Tailored phishing, malware variation, vulnerability research, and exploitation support all become cheaper and faster. That means more attackers can attempt more operations against more targets. It also means one-off attacks against smaller or previously uninteresting targets become economically viable.
In April 2026, the UK AI Security Institute (AISI) published its evaluation of Claude Mythos Preview and reported that the model had, for the first time, completed a full 32-step corporate-network attack simulation end-to-end — from initial reconnaissance through full domain takeover — on a purpose-built cyber range. The scenario is estimated to take a human professional around twenty hours. Mythos Preview solved it start-to-finish in 3 of its 10 attempts and, across all attempts, completed an average of 22 of the 32 steps. On the expert-level individual tasks that no model could finish before April 2025, the same model succeeded 73 percent of the time. Two years earlier, frontier systems could barely complete beginner-level capture-the-flag challenges.
The chart below shows how far each frontier model gets through a multi-stage attack chain — from M1 (initial reconnaissance) up through M9 (full network takeover) — as compute budget grows. The headline is simple: Mythos is the only model that finishes the chain. Every other frontier model, including Claude Opus 4.6 and GPT-5.4, stalls somewhere in the middle — credential theft, lateral movement, infrastructure compromise — no matter how many tokens you give it. Mythos keeps going and takes over the network. There is no visible plateau in the trend, and the newest model is the one pushing autonomous exploitation from partial to end-to-end.

The Institute was careful about what this does and does not mean. The cyber ranges lacked active defenders and defensive tooling, and the model paid no penalty for noisy actions that would have lit up a real security operations center. So this is not yet evidence that AI can autonomously compromise mature enterprises. It is evidence that AI can now autonomously compromise small, weakly defended networks once access has been obtained — which happens to describe a very large fraction of the real internet. The practical implication is that the floor of “too small to be worth targeting” is rising toward the ceiling. Attacks that previously would not have cleared the economic threshold of a human operator’s time now clear the threshold of a model’s runtime.
On defense, AI can eventually help close the gap, but institutions do not adopt new defensive capabilities at machine speed. They adopt through approvals, security reviews, procurement cycles, legal caution, and operational conservatism. Attackers do not have those constraints. They do not need to justify false positives to oversight bodies or worry about accidentally disrupting their own production networks.
That is the core near-term problem: AI improves both offense and defense, but offense can usually operationalize new capability faster.
The result is a dangerous middle period — let’s call it “the period of excessive badness“. Attackers gain speed before defenders gain reliable automation. Vulnerability discovery accelerates before patching becomes autonomous. Attack campaigns become cheaper before institutions redesign authorities and workflows for machine-speed defense.
That is why the next few years are likely to be harder, not easier.

Offense spreads first. Autonomous exploitation collapses the marginal cost of finding and weaponizing a vulnerability, while AI-scale code production floods the stack with insecure software faster than humans ever shipped secure software. Defense has no symmetric productivity gain ready to deploy: formal verification, memory-safe substrates, and AI-native detection all exist as research or point tools, not as infrastructure. The result is a period of excessive badness — years in which the hard-to-reach parts of the digital world (OT, embedded systems, medical devices, legacy enterprise, firmware) pay the steepest price, because they cannot be patched on the cadence that industrialized offense now operates at. E-commerce-grade trust degrades, not uniformly, but in the places where the rebuild wave hasn’t reached.
Then comes the rebuild — paid for by the losses of the bad years, the way every prior security regime was paid for by its own catastrophe. Proof-carrying construction becomes the default for new code. Verification moves from a research artifact into the build pipeline. AI red-teaming runs continuously against every service. Hardware roots of trust, memory safety, and wide-scope AI defense get deployed as infrastructure, not features. What emerges is a new cyber parity that looks like today’s in shape — commerce works, banks work, the internet functions — but sits on a different substrate. The equilibrium holds not because attackers got worse, but because the floor got rebuilt. The policy question of this decade is not whether that end state arrives; it is how wide the gap is allowed to grow in the meantime, and which systems are sacrificed to the delay.
What This Means
The real cyber debate is usually framed in the wrong terms. People ask which vendor wins, which framework matters, or whether AI will replace analysts. Those are secondary. The first-order question is whether defenders can structurally attack the asymmetries that make offense easier — what tech is needed to correct this balance and how much damage accumulates before they do.
These are the questions I’m asking:
- Can we compress the patch gap fast enough that working exploits do not enjoy a months-long free run?
- Can we reduce the practical burden of defending millions of doors?
- Can we make inevitable failures smaller, shorter, and cheaper?
- Can we automate enough of defense that institutions are no longer asking humans to operate at a speed and scale humans were never built for?
The AISI evaluation of Mythos Preview is useful mostly because it puts a clock on those questions. Two years ago, frontier models could barely complete beginner tasks. Today, one of them can solve an end-to-end corporate attack simulation that costs a skilled human most of a workweek. The next generation will not be worse at it. Every month between here and a rebuilt substrate is a month the excessive-badness period gets a little wider — and a few more of the hard-to-reach systems (OT, embedded, medical devices, legacy enterprise, firmware) get quietly written off as unsalvageable.
The end state is plausible. Proof-carrying code, memory-safe foundations, continuous AI red-teaming, wide-scope AI defense running as infrastructure — these are research artifacts and point tools today, and they become substrate in the decade ahead. What is not yet decided is how wide the gap gets on the way there, and who pays for it. That is a policy question, not a technical one.
DEFCON may never look like the bolt convention I argued it should. But the realistic win isn’t boring — it’s bounded. A period of excessive badness that ends. A rebuild that shows up in time. A new parity we actually reach, on a different substrate than the one that got us here. That, more than any single breach or product cycle, is the real story of where cyber is headed.ns from abstractions into a timer. Two years ago, frontier models could barely complete beginner tasks. Today, one of them can solve an end-to-end corporate attack simulation that costs a skilled human most of a workweek. The next generation is unlikely to be worse at it. The phases above are not a leisurely roadmap. They are the order in which defenders have to show results before the offense-defense gap becomes structurally self-reinforcing.
If the answer to those questions is no, cyber will continue to feel like a field where defenders work harder each year only to hold less ground. If the answer is yes, the next decade could mark a genuine transition from artisanal defense to industrial defense.
That, more than any single breach or product cycle, is the real story of where cyber is headed.
Sources:
- Our evaluation of Claude Mythos Preview’s cyber capabilities — AI Security Institute
- Claude Mythos Preview completes full cyberattack simulation for the first time — The New Stack
- Claude Mythos Preview shows “unprecedented” attack capability, warns AI Safety Institute — Computing
- Testing reveals Claude Mythos’s offensive capabilities and limits — Help Net Security
