softmax

Streamlining softmax by reducing inputs without sacrificing accuracy

Reducing softmax's inputs from N to N-1 is a neat theoretical insight, those redundant parameters are doing nothing.

4 min readMachine Learning

There is a quiet elegance in the observation that softmax, a function with N inputs and N outputs, actually only needs N-1 degrees of freedom. The user who posted this reasoning on our community board has spotted a redundancy baked into one of the most common operations in modern machine learning. Because the outputs must sum to one, one logit is always determined by the others. Enforcing a constraint, say, that the logits sum to zero, and calculating the last logit as the negative sum of the rest is mathematically sound. It removes a handful of parameters from the final layer. The question is whether that theoretical neatness translates into a practical win.

We respect the impulse to trim the unnecessary. It aligns with the kind of thinking we see in work like Solving functional gradient descent with adaptive representations, where the goal is to make representations more efficient without sacrificing what they can express. But here, the benefits are almost certainly negligible for most practitioners. The number of parameters removed is exactly the number of classes minus one, tiny relative to the rest of a model. Convergence speed might improve, but as the author admits, that intuition could be wrong. And there is a real cost: breaking from the standard softmax implementation means your model is no longer a drop-in replacement for every pretrained checkpoint or library expectation. You gain a sliver of theoretical purity and lose compatibility.

What would we tell a reader who asked whether to implement this? Start by asking what problem you are actually solving. If you are training a small model on a task with only a few classes and you have hit a wall on overfitting, this constraint could act as a mild regularizer worth testing. But for the vast majority of use cases, the practical takeaway is this: the degrees-of-freedom argument is intellectually satisfying, but the overhead of deviating from standard practice outweighs the marginal gain. This is a reminder that not every mathematically valid simplification is a useful engineering choice. Compare this to the approach in Transform Your Style: Google Photos AI Wardrobe Now Accessible on All Devices, where an AI feature was deliberately broadened to reach more users. That decision prioritized accessibility over a more constrained, theoretically cleaner design. The same principle applies here: sometimes the best solution is the one that works everywhere, not the one that is perfectly minimal.

The most interesting open question is whether this constraint could be exploited during training in a way that is invisible at inference time. If you can train with N-1 logits and then pad back to N before deployment, you might get the regularization benefit without the compatibility cost. Until someone tests that, the answer to the original question, is there a good reason not to do this?, remains a pragmatic yes: the benefit is too small to justify the friction.

From Machine Learning

Softmax has N inputs and N outputs but it's output only has N-1 degrees of freedom because of the condition that the sum of outputs must be equal to one. Based on this we can figure out that actually we can make due with only N-1 inputs by making an assumption that logits before softmax must sum up to zero (though it can be any other constant value) and have the last logit be calculated as minus sum of all the other logits. In theory it should remove "unnecessary" parameters from the last layer before softmax (however few of them…

Read the original at Machine Learning