The question posed by a machine learning practitioner on Reddit cuts to a core tension in classification work: when do you admit that your data is too sparse to support the distinctions you are asking a model to learn? Grouping rare dog breeds into a catch-all "Other" category feels like a pragmatic fix, but as the original poster suspects, it may create more problems than it solves. Their intuition about weirdly-shaped hyperplanes is not just technically plausible; it points to a deeper issue of whether a model can meaningfully separate classes that are defined by exclusion rather than shared characteristics.
We have seen similar trade-offs play out in adjacent fields. For instance, when exploring real-world computer vision, the pressure to maximize accuracy on a long tail of rare classes often leads to brittle systems that fail on the very diversity they were meant to handle. The "Other" category becomes a dumping ground where the model learns to place anything that does not fit neatly elsewhere, which is a recipe for inconsistent predictions. And while the original poster is right to consider out-of-distribution detection as an alternative, that approach brings its own complications, particularly when the boundary between "known" and "unknown" is fuzzy. This is not unlike the challenge of clean data starts with catching AI slop, where filtering out noisy samples can inadvertently remove signal that the model needs to generalize.
Our honest take is that grouping classes is a decision that should be driven by the downstream use case, not just by sample counts. If the goal is to distinguish between a handful of common breeds, then collapsing rare ones into "Other" is a reasonable way to simplify the problem. But if the model will encounter those rare breeds in production, then throwing away their samples is a form of self-deception. The model will not learn to say "I do not know" unless you explicitly train it to do so, and a catch-all category does not accomplish that. It just teaches the model that there is a single, ill-defined blob of "everything else" that it can ignore. The better question is whether you can tolerate the model being wrong about those rare cases, or whether you need to invest in collecting more data for them, even if that means accepting lower overall accuracy.
What we would tell a reader who asked us about this is straightforward: do not treat grouping as a default move. Treat it as a deliberate modeling choice with clear consequences. If you do group, consider whether the "Other" category should be a separate, one-vs-rest classifier rather than a single label in a softmax layer. And if you are tempted to drop the rare classes entirely, remember that you are also dropping the ability to detect them as anomalies. For a practical takeaway, consider this: the shape of your decision boundaries is a reflection of the questions you ask. If you ask a model to separate chihuahuas from wolf-like dogs within an "Other" category, you are asking it to compress a huge amount of visual variety into a single point of similarity. That is a losing bet. Instead, let the model focus on what it can learn well, and treat the rest as a separate problem. The open question is whether you are willing to accept the operational complexity that comes with that honesty.