The hidden risk in AI plant identification: closed-set models can't say "I don't know.

In the pursuit of safe and accurate plant and fungi identification, I abandoned YOLO due to its closed-set classification limitations.

3 min readMachine Learning

A model that cannot say "I don't know" has no place in a tool meant to keep people alive. That is the blunt lesson from one developer's work on a handheld device for identifying wild plants and fungi. He trained specialist YOLO models on iNaturalist data, hit 94 to 96 percent accuracy across his target species, and then discovered the gap that accuracy numbers hide. Feed a closed-set model an image it has never seen, and it will still pick one of its known classes with near total confidence. In foraging, that is not a bug. It is a lethal design flaw.

The developer's fix is worth studying, not because everyone needs to build a Hailo 8L pipeline tomorrow, but because it exposes a deeper truth about how we should think about AI in safety-critical roles. Softmax confidence, the number most people treat as a measure of certainty, is normalized across a closed set. There is no probability mass left over for "none of the above." Confidence thresholding fails because out-of-distribution images produce scores indistinguishable from real ones. His layered response, energy scoring on raw logits, ensemble disagreement, and a retrained "none of the above" class, treats the problem as structural rather than cosmetic. That is the right instinct. You do not tune your way out of a category error. You change the architecture.

What matters here is not the specific model names or the 13 TOPS compute budget. What matters is the principle: in any application where a wrong answer carries real cost, the system must be built to refuse. A router that rejects before classification, an energy score that separates known from unknown, a model that can say "I don't know" and mean it. These are not optional extras. They are the difference between a tool that assists a human and one that confidently misleads them. The developer found that energy scoring outperformed native confidence thresholding by a wide margin. That result should push every team building safety-relevant AI to ask whether their own evaluation metrics reward false certainty.

The practical takeaway is direct. If you are building a model for any domain where a wrong prediction can harm someone, stop optimizing for accuracy alone. Build rejection into the system from the start. Test with out-of-distribution data, not just held-out samples from the same distribution. Measure how often your model says "I don't know" when it should, and make that a first-class performance metric. The open-source community gets a concrete example of how to do this, and the rest of us get a reminder that confidence is not competence. The model that refuses to guess is the one worth trusting.

From Machine Learning

I’ve been building an open-sourced handheld device for field identification of edible and toxic plants wild plants, and fungi, running entirely on device. Early on I trained specialist YOLO models on iNaturalist research grade data and hit 94-96% accuracy across my target species. Felt great, until I discovered a problem I don’t see discussed enough on this sub.

Read the original at Machine Learning