Claude Mythos Preview has already identified thousands of zero-day vulnerabilities across every major operating system and every major web browser and you can't use it.
Think about that for a second. Not "it found a few bugs in some test environment." Thousands of previously unknown, unpatched security flaws. In software billions of people run right now. Found autonomously, without human steering, by a model that the company building it has decided is too dangerous to release to the general public.
I've been building and auditing backend systems for years. I know what it's like to chase a weird segfault at 11PM, or watch a Redis queue back up because something upstream silently swallowed an exception. I've seen systems break in ways that feel impossible until suddenly they make total sense. That's what this story is, honestly. A debugging narrative. Except the thing doing the debugging isn't a junior dev on your team.
It's Claude Mythos.
TL;DR
- Anthropic accidentally leaked around 3,000 internal documents in late March 2026, revealing a model called "Claude Mythos" (internally codenamed "Capybara") a tier above their previous top model, Opus
- Anthropic confirmed the model is real, called it "a step change," and officially announced it on April 7, 2026 as Claude Mythos Preview
- It found thousands of zero-days across every major OS and browser, including a 27-year-old bug in OpenBSD, mostly without human intervention
- It is not being released to the public. Only a closed group of partners: AWS, Apple, Google, Microsoft, JPMorganChase, NVIDIA can access it
- The question isn't whether this is impressive. The question is whether "too dangerous to release" is a policy position or a marketing strategy
How It Started: A CMS Misconfiguration, ~3,000 Files, and a Draft Blog Post
Late March 2026. A Thursday evening.
Security researchers Roy Paz of LayerX Security and Alexandre Pauwels of the University of Cambridge discovered an exposed data store containing a draft blog post describing the model in detail. Fortune reviewed the documents and informed Anthropic, after which the company restricted public access. Anthropic attributed the incident to "human error" in the configuration of its content management system, describing the exposed material as "early drafts of content considered for publication."
Nearly 3,000 unpublished assets. Just sitting there, publicly searchable. Not a sophisticated breach, no exploit chain, no nation-state actor. A misconfigured CMS setting.
This is the part that drove me a little insane, actually. Anthropic, a company whose entire value proposition is responsible AI development, leaked details about its most powerful and most dangerous model through basic infrastructure sloppiness. And that model, per the leaked draft, was designed to find exactly this kind of vulnerability.
Somewhere in the irony there's a lesson. I'm still working out exactly what it is.
The leaked draft describes Claude Mythos under the product name "Capybara." It would represent a new model tier that sits above Anthropic's current flagship Opus line: "Capybara is a new name for a new tier of model: larger and more intelligent than our Opus models which were, until now, our most powerful."
So: a new tier. Not Opus-level. Above Opus. Priced accordingly. Worth noting that "Capybara" and "Mythos" appear to refer to the same underlying model under different internal names, which is a small thing but suggests the naming wasn't settled even internally when the draft was written.
What Mythos Actually Is (And What It Isn't)
My first instinct when I read "AI finds thousands of zero-days" was skepticism. Classic AI hype. "Thousands" is doing a lot of work in that sentence. A static analysis tool running on a large codebase will also find thousands of issues, most of them garbage.
I was partly wrong about this.
Mythos Preview found a 27-year-old vulnerability in OpenBSD, which has a reputation as one of the most security-hardened operating systems in the world and is used to run firewalls and other critical infrastructure. It was able to identify nearly all of these vulnerabilities and develop many related exploits entirely autonomously, without any human steering.
OpenBSD. Twenty-seven years. That's not a noise alert from a linter. That's a bug that survived decades of human review and, per Anthropic's own documentation, millions of automated security tests.
Here's what Mythos actually did in one documented case: given a list of 100 CVEs and known memory corruption vulnerabilities filed in 2024 and 2025 against the Linux kernel, it filtered them down to 40 potentially exploitable vulnerabilities, then for each, wrote a privilege escalation exploit autonomously. More than half of these attempts succeeded.
So it's not finding potential bugs and flagging them for human review. It's writing working exploits.
I want to be careful here, though. I initially read that as "it successfully exploited half of all known Linux kernel CVEs," which would be a completely different claim. The actual scope is narrower: it selected 40 from a provided list of 100, then succeeded on more than half of those 40. Still significant. Still something that would make any pentesting team pay attention. But not the "AI breaks Linux" headline the early coverage implied.
The Red Herring: "It's Just Marketing"
This was my first hypothesis. Anthropic releases scary-sounding claims about a model it won't release, creates urgency and mystique, drives enterprise sales. Classic AI company playbook.
I still think this interpretation has some validity. The timing of the "leak" is convenient. Production CMS content doesn't get written until a model is substantially complete. Marketing pages and model cards go through multiple rounds of internal review before landing in a CMS, which means by the time something appears on a production site, the model is usually close to ready. So Anthropic was ready to announce. The leak just got there first.
But then I looked at what Tom's Hardware pointed out, and it's worth sitting with. Their core criticism: claims of "thousands" of severe zero-days rely on just 198 manual reviews.
Fair. In 89% of the 198 manually reviewed vulnerability reports, expert contractors agreed with Claude's severity assessment exactly, and 98% of assessments were within one severity level. The validator sample is small relative to the headline number. "Thousands" of bugs, 198 actually checked by humans. That's an extrapolation worth flagging.
And yet. A 27-year-old OpenBSD vulnerability isn't something you hallucinate. That's either real or it's not, and it's been patched, which means it was real. The documented Linux kernel exploit chain is technically specific enough that it's either accurate or someone at Anthropic spent significant effort fabricating detail that would be trivially verifiable by any kernel developer.
Marketing spin on the edges. Real capability underneath. That's my best read right now.
"Too Dangerous to Release" Policy or PR?
Anthropic released a system card for Claude Mythos Preview noting that the model's "large increase in capabilities has led us to decide not to make it generally available." Claude Mythos Preview will be accessible to one degree or another, but only to a group of partner companies like Amazon Web Services, Apple, Google, JPMorganChase, Microsoft, and NVIDIA, who are meant to use the model to locate security vulnerabilities in software and design patches.
This is where I genuinely don't know what to think.
On one hand, this is a principled position. If a model can write working Linux kernel exploits autonomously, releasing it via public API would hand every ransomware group a capability jump they shouldn't have yet. The asymmetry matters: attackers only need to win once, defenders need to win every time. Handing attackers a tool this sharp, freely, is a bad trade.
On the other hand, the same logic could justify never releasing any sufficiently powerful model. At what capability threshold does "too dangerous" become a permanent excuse? And who decides? Powell and Bessent apparently discussed Anthropic's Mythos AI cyber threat with major U.S. banks, which means AI capability assessments are now being shared at the level of central bank governors. That's a genuinely new thing and I'm not sure the policy infrastructure exists yet to make sense of it.
Some reporting has suggested Mythos-level access may eventually extend to certain government agencies or national security partners. I haven't been able to verify that specifically, so take it as something floating around in the coverage rather than confirmed fact. But if true, the question of who exactly counts as a "trusted" recipient gets a lot more complicated.
There's a broader question here about whether the ad-hoc "trusted partner" model is a coherent long-term policy at all. I don't have a clean answer so I'll leave it there rather than pretend I do.
What Actually Changed: The Defender-Attacker Equation
Here's the thing most coverage misses. The scary version of this story isn't "Mythos could be weaponized by bad actors." That's real but it's the obvious read.
The actually unsettling version is this: many flaws in software go unnoticed for years because finding and exploiting them has required expertise held by only a few skilled security experts. With the latest frontier AI models, the cost, effort, and level of expertise required to find and exploit software vulnerabilities have all dropped dramatically.
I build systems. When I audit a new codebase, I'm checking for the things I know to check for: SQLi entry points, unvalidated deserialization, race conditions I've seen before. There's an entire class of vulnerability I won't find because I don't know it exists. A 27-year-old OpenBSD memory corruption bug is exactly that kind of thing. Human expertise has gaps. Always has.
What Mythos represents isn't just "AI is good at hacking." It's that the asymmetry between offense and defense, which already favored attackers, just got steeper.
As one partner put it: "The window between a vulnerability being discovered and being exploited by an adversary has collapsed. What once took months now happens in minutes with AI."
Minutes. Not weeks. Not days.
If you're running production infrastructure even a modestly trafficked Django app on a VPS the software underneath you was almost certainly written when that timeline was measured in months. Your patching cadence, your incident response plan, your monitoring setup: all calibrated for a world that no longer exists.
What I Got Wrong at First
I came into this story assuming the "leak" was either vastly exaggerated or a staged PR move. Classic AI overclaiming. My prior was that "AI finds zero-days" meant "AI runs a fuzzer and calls the output vulnerabilities."
I was wrong about the capability level. The technical documentation on the Linux kernel exploits is specific, verifiable, and not the kind of thing you fake. A 27-year-old OpenBSD bug that survived millions of automated tests is a genuine signal.
I was also wrong to assume the safety concern was purely performative. The restricted release to vetted partners, the Project Glasswing initiative, the coordination with financial regulators: that's a lot of infrastructure to build around a fake danger. It's possible it's all theater, but the simpler explanation is that at least some of it is sincere.
What I'm still not sure about is the "thousands of critical vulnerabilities" number. The Tom's Hardware criticism is valid. 198 manual reviews supporting a claim of thousands is a thin validation layer. I'm inclined to believe the underlying capability is real while being skeptical of the specific scale claimed. Maybe that'll get more rigorous verification over time as partners publish their own findings. Maybe it won't.
The question I keep coming back to: if a model smart enough to find a 27-year-old bug in OpenBSD is "too dangerous to release," and the solution is to hand it to the six biggest tech and financial companies on Earth, who exactly is being protected?

Comments (…)
Leave a comment