softmax

Streamlining softmax by reducing inputs without sacrificing accuracy

Reducing softmax's inputs from N to N-1 is a neat theoretical insight, those redundant parameters are doing nothing.

4 min readMachine Learning
From Machine Learning

Softmax has N inputs and N outputs but it's output only has N-1 degrees of freedom because of the condition that the sum of outputs must be equal to one. Based on this we can figure out that actually we can make due with only N-1 inputs by making an assumption that logits before softmax must sum up to zero (though it can be any other constant value) and have the last logit be calculated as minus sum of all the other logits. In theory it should remove "unnecessary" parameters from the last layer before softmax (however few of them…

Read the original at Machine Learning