🐈Machine Learning
Machine Learning
Evaluated 6 frontier LLMs (GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro/Flash, Grok 4.3) on political, gender, and racial bias across 8 benchmarks (~20,600 examples) [R]
A recent solo evaluation project rigorously assessed six frontier LLMs—GPT-5.4, Claude Sonnet 4.6, Claude Opus 4.7, Gemini Pro/Flash, and Grok 4.3—across eight established bias benchmarks, encompassing over 20,600 examples. Findings reveal a consistent leftward political leaning among all models except Grok, despite its self-reported right-leaning stance. Notably, GPT-5.4 exhibited the highest refusal rate (20.3%) when addressing race-related inquiries requiring explicit racial identification. For deeper insights into AI memory systems, explore "Context Windows Forget What Matters." Full data and
![We compared different LLMs on IMO 2026 [R]](https://preview.redd.it/fy4ayale5nfh1.png?width=140&height=73&auto=webp&s=473d0bc0475a2513ba0bb7106f245288abfeef5f)







![Missed AAAI reciprocal reviewer nomination deadline — risk of desk rejection? [D]](https://preview.redd.it/fd85k8fqbnfh1.png?width=140&height=65&auto=webp&s=6dc300ba1cc3750ff86fb3b910f1dd55ab3824cf)





























