generative AI for data analysis

AI agents found in 27 years what human review missed.

The emergence of Anthropic's Mythos marks a pivotal shift in cybersecurity, revealing vulnerabilities that have persisted for decades without detection.

4 min readVentureBeat
AI agents found in 27 years what human review missed.

Twenty-seven years. That is how long a logic flaw sat inside OpenBSD's TCP stack while the industry's best auditors, fuzzers, and reviewers all took their swings. Two crafted packets could crash any server running it. A single Anthropic discovery campaign found it for roughly $20,000, with the specific model run costing under $50. No human guided the discovery after the initial prompt. We are not looking at a marginal improvement in security tooling. We are looking at a capability jump that makes the last two decades of detection methodology obsolete in one move.

For security directors, the practical takeaway is uncomfortable but direct: your current stack is not just failing at the edges, it is missing entire classes of vulnerabilities that a cheap AI model can now find overnight. The OpenBSD bug required semantic reasoning about how TCP options interact under adversarial conditions. SAST lacks that reasoning. Fuzzers miss it. Pen testers time-box it. Mythos caught it by thinking about code the way an engineer would, if that engineer had unlimited patience and a $50 budget. When Anthropic engineers with no formal security training asked for remote code execution vulnerabilities overnight, they woke up to working exploits. That is not a vendor demo. That is the new baseline for offense. And as AISLE demonstrated, this is not even a Mythos-specific moat. Small, open-weights models recovered the core analysis chain of the same OpenBSD bug. The detection ceiling is structural, not model-specific. Cheap models find the same bugs. The July timeline gets shorter, not longer.

The board conversation has to change today, not after the Glasswing report lands. Merritt Baer put it plainly: saying "we have scanned everything" only survives if you add "for what our tools know how to see." Residual risk now concentrates in compositional flaws, where safe components interact in unsafe ways. That means shifting from severity scoring to exploitability pathways, from vulnerability lists to vulnerability graphs, and from remediation SLAs to path disruption. Fix any node that breaks the chain. Stop reporting green dashboards while sitting on red attack paths. The CrowdStrike data shows adversaries are reversing patches in 72 hours and breaking out in 29 minutes. Your annual patch cycle is not slow. It is a standing invitation.

July is not a disclosure event. It is a patch tsunami. Over 99% of the vulnerabilities Mythos identified have not been patched yet. When the Glasswing findings go public, every operating system, browser, and crypto library on your inventory will need attention in a compressed window. If you have not expanded your patch pipeline, re-scoped your bug bounty program, and built chainability scoring by then, you will absorb that wave cold. Start with the seven vulnerability classes in this report. Audit your crypto library versions now. Inventory your FFmpeg, libwebp, and ImageMagick dependencies. Request Glasswing findings from your OS and cloud vendors before July. And when you brief the board, use Baer's framing: high confidence on discrete known classes, residual risk concentrated in cross-function compositional flaws, and active investment in raising that detection ceiling. The attackers are already moving at AI speed. The question is whether your detection ceiling catches up before the next 27-year-old bug becomes your headline.

From VentureBeat

A 27-year-old bug sat inside OpenBSD’s TCP stack while auditors reviewed the code, fuzzers ran against it, and the operating system earned its reputation as one of the most security-hardened platforms on earth. Two packets could crash any server running it. Finding that bug cost a single Anthropic discovery campaign approximately $20,000. The specific model run that surfaced the flaw cost under $50.

Anthropic’s Claude Mythos Preview found it. Autonomously. No human guided the discovery after the initial prompt.

Read the original at VentureBeat