The moment you start retraining on your own model's inferences, you're no longer just correcting a classifier; you're negotiating with its confidence. The user who posted this is right to be uneasy. Their boss's suggestion to fold model outputs back into the training set, even with human corrections, is a recipe for silent drift. The real problem isn't the retraining frequency, it's the assumption that more data automatically means better data. With six thousand rows, a dominant class, and a single lonely example at the bottom, the margin for error is already razor thin.
A sliding window of recent data sounds agile, but it's a trap when labels disappear for long stretches. If the model stops seeing a rare class for three months, the window forgets it, and the next time that label reappears, the classifier has to relearn it from scratch. That's not adaptation; that's amnesia. A replay buffer, a small curated set of examples for rare or historically important classes, is the smarter anchor. It preserves institutional memory without letting stale data bloat the training set. Pair that with a recent window for the dominant, fast-moving patterns, and you get a policy that respects both novelty and continuity.
On the question of incremental learning versus full retraining, the honest answer is that partial_fit and stream-learning libraries are seductive but fragile here. They assume a stable label space, and this problem has the opposite. Periodic full retraining, done on a disciplined schedule, gives you a chance to audit what changed, why it changed, and whether the model's own inferences are quietly reinforcing a bias toward the majority class. That audit is not optional. You need a golden test set, fixed and untouched by any retraining loop, to catch collapse before it becomes a production incident. Rolling tests are useful for monitoring drift, but they can't tell you if the model has lost the plot; only a fixed benchmark can do that.
Evaluate per class, always. Accuracy on a dataset where one class owns eighty percent of the labels is a lie. Track precision and recall for the smallest classes as if they were the only ones that mattered, because to the person who needs that rare label, they are. And when a new label appears, don't rush to add it. Wait until you have enough human-verified examples to train on, even if that means launching without it for a cycle. The user's instinct to question the boss's plan is sound. The fix is not more retraining; it's more deliberate retraining, with memory, with measurement, and with the patience to let rare classes earn their place.