vulnerability scanning

Explore how an open-source AI agent refines vulnerability detection.

False positives have long made AI-powered code scanning feel like more trouble than it's worth.

3 min readInfoQ
Explore how an open-source AI agent refines vulnerability detection.

Google's decision to open-source Mantis lands at a telling moment. The tool is an AI-agent framework built to automate the entire vulnerability lifecycle: finding issues, validating them, reproducing them, and even patching them. The stated motivation is just as important as the technology itself. Google says Mantis exists to counter the high rate of false positives and hallucinated vulnerabilities that plague conventional AI-powered code scanning. That admission is the real story here, because it signals a maturing of the industry's relationship with generative AI. We are no longer celebrating raw output. We are demanding that output be trustworthy.

For anyone who has spent time inside a modern security dashboard, the pain point is familiar. AI scanners can flag a suspicious code path with confidence, only for a human to spend twenty minutes confirming the issue is a phantom. The damage is not just wasted time. It is the erosion of trust in the very tools meant to protect us. Mantis aims to close that loop by making the agent responsible for the whole arc of a vulnerability, not just the initial detection. That is a meaningful step forward. It moves the conversation from "AI can find things" to "AI can be held accountable for what it finds." In that sense, Mantis is less about a clever new feature and more about enforcing a standard of rigor that has been missing.

The timing also aligns with broader concerns about agentic systems acting without proper guardrails. We recently saw AI agents shared user images, highlighting data security concerns, a reminder that autonomy without oversight can create new attack surfaces. Mantis does not eliminate that risk, but it does narrow the blast radius by focusing on code-level verification rather than broad, unconstrained behavior. Similarly, the pressure on security teams has never been higher, especially as threats become more brazen. The North Korean hackers linked to $351M Bitget crypto theft shows how quickly a single overlooked flaw can be weaponized at scale. Mantis will not stop determined adversaries by itself, but reducing the noise in vulnerability reports gives defenders a better shot at spotting the real signals before they become incidents.

Our take is straightforward. Open-sourcing Mantis is the right call, not because it will instantly secure every codebase, but because it forces the AI security space to grow up. The benchmark for these tools should not be how many alerts they generate. It should be how many they save us from chasing. Google is effectively saying that a scanner that cries wolf too often is worse than one that stays silent. That is a mature position, and it sets a useful precedent for the rest of the industry. The specific detail to watch is how the community handles the validation step. If Mantis can reliably demonstrate that a reported vulnerability is real and exploitable, it will change how we budget for security work. If it merely moves the hallucination problem from detection to confirmation, then we have only traded one headache for another. For now, the promise is real, and the direction is right. The next twelve months will tell us whether the agents are ready to do the boring, exacting work that security actually requires.

From InfoQ

Google has open-sourced Mantis, an AI-agent framework designed to automate the software vulnerability lifecycle, from identifying and validating vulnerabilities to reproducing and fixing them. Google says it developed Mantis to address the high rate of false positives and hallucinated vulnerabilities produced by conventional AI-powered code scanning.

Read the original at InfoQ