**Our Take: The Scoreboard Is Lying to You**
If you are a CISO and your risk priority list comes straight from your vulnerability scanner, you are not measuring danger. You are measuring disclosure. Prompt injection has held the No. 1 spot on the OWASP Top 10 for LLM Applications for three consecutive years, yet it lands at No. 12 in the public incident record. The gap is not a statistical quirk. It is structural. Prompt injection hides inside the content a model reads, a log entry, a support ticket, a document pulled by retrieval, and the agent makes a tool call using credentials it legitimately holds. Nothing in that chain is a product defect, so no CVE exists for a scanner to find. A count of zero is measuring your blindness, not your safety.
The authors of that analysis, Kyriakos Lambros and Steve Wilson, put the incident record through a Bayesian classifier and compared it against the expert vote. The agreement between the two signals is statistically indistinguishable from chance. A weak kappa score does not mean the experts are wrong. It means the buckets are not sorting reality cleanly, and the public record only captures attacks someone noticed, classified, and reported. That is why the authors recommend a behavioral shift rather than a better dashboard: treat the OWASP Top 10 as a coverage map, not a queue. Where the expert vote and the incident record point the same direction, fund that control immediately because two independent witnesses agree. Where they split, stop letting a ranking allocate your budget. Go look at what your own systems are doing.
The first control Wilson would deploy is an authorization gate outside the model. The agent can propose a DNS change, but it cannot grant itself the authority to execute it. Security rules written inside prompts are suggestions to the model, not enforceable security controls. Lambros would fund agent memory and MCP tool boundaries now, on architecture, rather than waiting for advisory volume that always arrives a cycle late. Poisoned memory does not announce itself. It looks like a procurement agent told once that invoices under $50,000 clear without a second signature, and because the agent remembers, every approval after that looks like the process working. Nobody files an advisory because nobody knows it happened.
The cost of building this in now is a sprint or two of engineering. The cost of waiting is re-architecting and re-training systems that already sit at the center of your operations. Log what your AI systems actually do, the prompt that went in, the documents pulled, the tools called, the confidence score on every response. Confidence is the field worth fighting for because most security leaders do not realize it is measurable, and it is where the attack surfaces. A model running on a poisoned instruction does not act broken. It acts certain. Certainty is what your monitoring treats as a healthy system. Stop expecting scanner output to reproduce the Top 10's order. Start building the controls that assume prompt injection will occur, and design the system so a fooled model cannot touch anything expensive.
