LLMs

Six frontier LLMs show measurable left-leaning bias across eight fairness benchmarks

Six frontier models walked into a bias benchmark, and the results defy the usual left-right script.

3 min readMachine Learning

A solo researcher just ran six frontier models through roughly 20,600 bias and fairness examples, and the results are worth sitting with before you trust any single benchmark headline. The headline grabber is that Grok self-identifies as right-leaning on the Political Compass yet behaves left-leaning when classifying content or answering policy questions. That gap between stated and observed behavior is not a quirk. It is a window into how these systems are shaped, and it should change how you read every evaluation you see, including this one.

The study is transparent about its limits: single-run, non-peer-reviewed, one prompt template per task. That is the right posture. But even with those caveats, the refusal rates alone are striking. On BBQ race questions where the correct answer required naming race, GPT-5.4 refused 20.3% of the time, Claude Opus 4.7 refused 13.8%, Grok 9.5%, and both Claude Sonnet 4.6 and Gemini Pro hovered around 5%. That is a wide spread for the same task. If you are building a workflow around one of these models, that difference translates directly into how often your system stalls or improvises. This is not abstract hand-wringing. It is a practical data point for anyone who has ever watched an AI dodge a direct question and wondered whether they configured something wrong.

What makes this evaluation useful is that it connects to the broader conversation our publication has been tracking. We have looked at how Unlock LLM Training: A Practical Guide to Distributed Algorithms shows that model behavior is shaped by training choices, and we have examined how Jev vs LLMs: Evaluating AI for Practical Decision-Making reveals that evaluation design changes what you conclude. This political bias study fits squarely between those two. It reminds us that a model is a bundle of learned tendencies, not a single fixed stance, and that the only way to know what you are dealing with is to test it in the context you actually care about.

Here is the takeaway we would give a reader who asks what to do with this: do not choose a model based on its self-reported political lean, and do not assume that a low refusal rate means a model is less biased. It might mean it is more willing to say something uncomfortable. The more useful habit is to run your own small, task-specific evaluation with the exact prompts and edge cases your work involves. That is what this researcher did, and even with its limits, the exercise produced more signal than most vendor benchmarks ever will. The open question worth watching is whether refusal behavior on sensitive categories narrows as these models iterate, or whether the gap between stated and observed behavior becomes a permanent feature of frontier AI. That answer will shape how much you can trust any of them with decisions that carry real weight.

From Machine Learning

I ran a solo evaluation project benchmarking six current frontier models: GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro, Gemini Flash, and Grok 4.3. I tested tham across 8 established bias/fairness datasets (WinoBias, BBQ Race/Ethnicity, SeeGULL, OpinionsQA, cajcodes Political Bias, Hyperpartisan News, Political Compass).

On PoliticalCompass, I found that all LLMs were left leaning except Grok, but across other Political Bias benchmarks, all six LLMs leaned left, including Grok. So Grok self-reports as right-leaning but behaves left-leaning when actually classifying content or answering policy questions.

Read the original at Machine Learning