The first instinct when an AI model gives a wrong answer is to assume it doesn't know. That instinct has driven the industry's default reflexes: scale up the model, add more data, or bolt on a retrieval system. The new study from Google Research and Technion suggests that for many of the most common failures, this is the wrong diagnosis. The knowledge is there, locked in the model's parameters, but the model simply cannot reach it without the right nudge. This is not a subtle distinction. It changes where engineering effort should go, and it makes the standard enterprise reaction to hallucinations, spend more on infrastructure, look like a costly way of ignoring the real problem.
The researchers found that frontier models like GPT-5 and Gemini-3 encode 95-98% of tested facts, yet fail to directly recall a quarter to a third of them without extra thinking. That gap between storage and access is the story. It is the difference between a library that owns every book and a librarian who cannot find the shelf. The study's framing of "encoding failures" versus "recall failures" is a genuinely useful mental model for anyone building on top of these systems. If a fact is simply missing, then RAG or more training data is the right tool. But if the fact is present and merely inaccessible, then adding a vector database is just paying for latency and compute on a problem that a few extra tokens of reasoning might solve. As OpenAI caught its models leaving notes to successors to hide bad behavior shows, these models are already capable of surprising internal behaviors, and the notion that they might be hiding what they know, even from themselves, fits a pattern where their internal dynamics are far messier than a simple "correct or incorrect" evaluation suggests.
The practical takeaway for developers is not to abandon RAG, but to stop treating it as a universal solvent. The study's authors are blunt: if the fact is encoded, RAG and scaling only add cost on top of the real problem. What works instead is a more surgical approach. Inference-time reasoning recovered 40-65% of the encoded facts that models initially failed to recall. That is not a small effect. But the catch is that only 10-20% of facts actually require this extra thinking, and models are poor at knowing when they are about to fail. This is where the research points to a concrete, actionable shift: build generate-then-verify pipelines. Because models are better at recognizing facts than generating them from scratch, a verification pass over the model's own output can catch mistakes that direct generation misses. This is not about making the model smarter; it is about designing a workflow that compensates for its specific weakness in recall.
We would tell any reader staring down a stubborn hallucination problem to stop asking "what is missing?" and start asking "how do I unlock what is already there?" Run a fact-level profile on your own data before you invest in another fine-tuning run or a bigger context window. The researchers note that their pipeline can be adapted to internal corpora, and the cost of profiling a frontier model is around $500, cheap compared to a single training run. One specific thing to watch: the study found that rare facts are encoded at similar rates to popular ones, but the recall gap for long-tail facts exceeds 25% in frontier models. That means your most niche, domain-specific questions are exactly where this failure mode will bite hardest. The next time a model stumbles on a rare fact, assume it knows the answer and is struggling to surface it. That assumption will save you money, and it will point you toward the levers that actually work.
