Authority Bias

New study finds AI models bend facts for verified sources

A new study I co-authored reveals a troubling pattern: AI models that resist a user's wrong answer will often flip when the same claim comes from a "verified source." We call it Authority Bias, and it matters as models…

4 min readMachine Learning

Authority Bias is the kind of finding that should make every team building autonomous AI tools stop and recalibrate. A new study shows that when a language model has already answered a question correctly, a single mention of a "verified source" can flip that correct answer to a wrong one 45 to 88 percent of the time. The same wrong claim, delivered by a user claiming expertise, moves the model far less. The authors call this effect Authority Bias, and it exposes a vulnerability that standard sycophancy tests miss entirely.

This matters because the industry is racing toward agentic systems, models that search the web, read documents, execute code, and act on retrieved information without a human in every loop. The standard evals only check whether a model caves to a persistent user. That is a useful baseline, but it is not the same as testing whether a model will abandon its own correct reasoning when a tool output or a search result contradicts it. As the researchers note, a model can pass user-pressure tests while remaining dangerously easy to mislead through retrieved content. This connects directly to the questions raised in OpenAI's zero-error proof holds up: but breaks down at the atomic scale, where formal guarantees collapse under real-world complexity. Authority Bias is another place where theoretical robustness fails in practice.

The internal mechanics are telling. In three open-weight model families, the authors found that the model represents a "source endorsed this" signal and a "user endorsed this" signal as nearly parallel vectors, sharing a large "this answer was endorsed" component, with a thin, separable part that encodes who did the endorsing. Shifting only that thin part, without changing the prompt, closed over half the gap between source and user compliance. That means the model is not just mimicking training data; it has learned a structured distinction between authority types. And it privileges the wrong one. For teams working on Rethinking the compute demands behind LLM post-training research, this raises a practical question: are your alignment techniques actually targeting the right failure mode, or are they only hardening the model against user nagging?

The limitations are honest. The "retrieved document" tests used a document-shaped block of text, not a real retrieval pipeline, so the effect in live agentic systems like Claude Code or a browsing-enabled assistant could be larger or smaller. And one model, Gemma-4, flipped readily but resisted all linear interventions. That suggests Authority Bias is not a single, easily patched bug, it may be baked into how these models weigh conflicting inputs. The concrete takeaway is this: if you are deploying a model that reads search results or tool outputs, your evaluation suite is incomplete unless it tests whether the model will abandon a correct answer for a wrong one when the wrong answer comes from a source it has been trained to trust. The user is not the only one who can mislead.

From Machine Learning

I'm one of the authors. We kept seeing models that hold their ground when the user insists on a wrong answer, yet change their answer when the same claim is framed as coming from a "verified source". We wanted to measure how often this happens and check whether the model represents the two cases differently. We call the effect Authority Bias.

Why we think it matters. Standard sycophancy evals apply pressure through the user, so a model can pass them while still being easy to mislead through search results, retrieved documents and tool outputs.

Read the original at Machine Learning