There is a right answer here, and it's the one that treats a Siamese network as a single model with shared weights, not as two separate networks that happen to be averaged. The original paper is undeniably thin on implementation details, which is why so many practitioners end up confused. But the core principle is simple: the same network processes two inputs, and the loss is computed on the pair, not on each input independently. That means the gradient flows through both branches of the forward pass, and the weight update is applied once, to the shared parameters.
The GitHub implementation the user found, where inputs are passed one after the other and the loss is computed for the last two, is not correct in spirit. It treats the network as if it were processing a sequence, which breaks the entire premise of a Siamese architecture. The whole point is that both inputs are evaluated by the same function at the same time, so the similarity or difference between their representations can be measured and backpropagated through both branches simultaneously. Passing them sequentially and then computing a loss on the final pair is a different model entirely, and it will not learn the kind of comparative representations the paper intends.
The second idea, using two copies of the network and then averaging the weights, is closer but still misses the mark. Averaging weights after every forward pass is a hack that introduces unnecessary noise and can destabilize training. The correct approach is to share the weights from the start, literally the same tensor, not two copies that get merged. This is what frameworks like PyTorch and TensorFlow make easy with their built-in weight-sharing mechanisms. You define the network once, pass both inputs through it, compute the loss on the two outputs, and call backward. The gradient is computed with respect to the shared weights, and the optimizer updates them once. That's it. No sequential passing, no averaging, no extra bookkeeping.
For anyone struggling with this, the practical takeaway is to stop thinking about "two networks" and start thinking about "one network, two inputs." The confusion is understandable because the architecture diagram in the paper shows two branches, but those branches are just the same function applied twice. Once you internalize that, the backpropagation becomes straightforward: you compute the loss on the pair, let the gradient flow through both branches, and update the shared weights in one step. If you're implementing this yourself, use a framework that supports weight sharing natively, and test it on a small dataset first to confirm the gradients are flowing as expected. That will save you hours of debugging and give you confidence that your implementation is faithful to the original intent.