ML models

Why stakeholders expect zero errors from models designed to improve decisions

There's a moment every data scientist knows well: the model performs safely in validation, stakeholders sign off, and then one wrong call triggers a cross-examination.

4 min readData Science

A data scientist builds a model that improves a key metric, clears every validation check, and gets a room full of stakeholders to nod along. Deployment day goes smoothly. Then, one wrong call. Suddenly, the same people who approved the model want to know why it failed. That frustration is not just about a technical misunderstanding. It is about a mismatch in expectations that no amount of model documentation will fix.

We have seen this pattern play out across the industry. Stakeholders are comfortable with the idea that humans make mistakes, but they treat models as if they should be held to a standard of perfection that even the most rigorous validation cannot guarantee. This is not a failure of communication. It is a failure of framing. When we present models as tools that "increase metric X at almost no cost," we skip over the part that matters: every model is a statistical approximation, and approximations have edges. The same way Unlock LLM Training: A Practical Guide to Distributed Algorithms breaks down complex systems into teachable components, we need to break down the reality of model error for stakeholders before the model goes live, not after.

The practical takeaway here is uncomfortable but necessary: if you cannot explain to a stakeholder why a model will occasionally be wrong, you have not finished building the model. You have only finished training it. That distinction matters because it changes how you present your work. Instead of leading with the accuracy number, lead with the error rate and what it means in context. Explain that a 95% accuracy rate on a validation set is not a failure to reach 100%. It is a ceiling, and it is a normal ceiling. This is where the conversation needs to shift from "will it work?" to "what happens when it does not?" The second question is the one that builds trust. The first one only builds hope, and hope is not a deployment strategy.

What would we tell a data scientist who is stuck in this loop? Stop treating stakeholder questions as interruptions and start treating them as part of the system. Build a small set of example failure cases into your presentation. Show them what the model gets wrong, and then show them what the cost of that error actually is. You will find that most stakeholders are not demanding perfection. They are demanding predictability. They want to know that the model fails in ways they can plan around, not in ways that blindside them. For a deeper look at how models navigate uncertainty at a structural level, Exploring Paragraph Structure: How LLMs Navigate Token Space offers a useful parallel: even the most advanced systems operate on probabilities, and understanding that is the first step to designing for it.

The next time a stakeholder asks why the model made a mistake, resist the urge to defend the model. Instead, ask them what they expected to happen. That question will reveal more about the gap between their mental model and yours than any chart ever could. The goal is not to eliminate the question. The goal is to make it a good question. And the only way to do that is to stop treating error as an anomaly and start treating it as a feature of the system that deserves as much attention as the accuracy metric. The model will keep making wrong calls. The question is whether the people around it will understand why.

From Data Science

Definitely the most frustrating thing as working as a Data Scientist. You run an experiment, find that building a model greatly increase metric X at almost no cost, has safe model metrics, present it to stakeholders, everybody agrees with proceeding to deploying and utilizing the model in production, and yet every time the model takes a wrong decision, we get questioned about it. Why did the model say this?

Read the original at Data Science