The promise of AI as an impartial judge is seductive. We are drawn to the idea of an evaluator that is fast, cheap, and endlessly scalable, especially when the alternatives are slow, expensive, and human. Bhaskarjit Sarmah's warning at the DHS 2026 workshop cuts to the heart of this allure, reminding us that these systems are not neutral arbiters but reflect the biases embedded in their training data. This is a critical moment to explore how AI agents learn by editing context, not model weights, because the mechanisms we use to build these tools directly influence the fairness of their judgments.
Our take is not that we should abandon LLM-based evaluation altogether. That would be throwing the baby out with the bathwater. The speed and scale they offer are genuinely transformative for tasks like initial content screening or flagging obvious errors. But the assumption that an LLM can objectively grade a nuanced research paper or fairly assess creative code is a dangerous overreach. We have seen how these models can inherit subtle cultural, linguistic, and even mathematical biases from their training. When you ask an LLM to judge, you are not getting an objective truth; you are getting a statistical reflection of the internet's collective, often flawed, perspective. This is why we believe the conversation must shift from "Can we trust LLMs?" to "How do we design evaluation frameworks that account for their limitations?" It is a practical matter of building guardrails, not a philosophical debate.
For our readers, this has immediate, actionable implications. If you are using LLMs to grade student work, you cannot simply assume the score is fair. If you are using them to rank papers, you must be aware that the model might favor verbose, confident-sounding text over genuinely novel ideas. The answer is not to return to purely manual processes, which are unsustainable at scale, but to use LLMs as a first-pass filter, not a final verdict. We would advise implementing a human-in-the-loop system where the AI provides a recommendation with a confidence score and a clear explanation of its reasoning. This is where understanding the underlying distributed algorithms used in LLM training becomes more than just technical trivia; it helps you understand why certain biases might be more pronounced. Similarly, a practical guide to using ChatGPT for work should include a critical eye on its outputs, not just a how-to on crafting prompts.
The most important takeaway is that an LLM is a tool, not a replacement for human discernment. It is a powerful accelerant for your work, but it is not a moral compass. The specific detail to watch for is the confidence calibration of the model. Does the LLM express high confidence when it is wrong? If so, that is a red flag. The next time you are tempted to automate a critical judgment call, ask yourself a simple question: Would you let a brilliant but biased intern make the final call without supervision? If the answer is no, then you should not let an LLM do it either. The future of evaluation is not about removing humans from the loop, but about making the loop smarter.
