Rethinking Attention: Distance-Based Scoring Puts Magnitude in Its Place

In a recent exploration of self-attention mechanisms, I replaced traditional dot-product attention with distance-based RBF-attention, prompting a reevaluation of how attention scores are calculated.

3 min readMachine Learning

There's a quiet bravery in asking what breaks when you swap out a core assumption, and that's exactly what this experiment does. Replacing the dot-product in self-attention with an RBF distance metric isn't just a clever hack; it's a direct challenge to a foundation the entire ML stack has quietly built itself around. The willingness to trace the domino effect, from instant memory blowups to the strange geometry of RoPE, shows that innovation isn't always about inventing something new. Sometimes it's about refusing to accept that the current way is the only way.

For practitioners, the practical takeaway is twofold. First, the math trick that turns a naive distance computation into a simple key-norm penalty is genuinely elegant. It means you don't need a fundamentally new architecture to test distance-based attention; you just need to adjust what you're optimizing. Second, the attention sink problem is a reminder that models will always find a way to game your objective. If magnitude bullying is how they manage uncertainty, then forcing a Euclidean geometry means you must provide a safe harbor, those zero-initialized register tokens are a clever, human-centered solution to a problem the model didn't ask for but needed solved.

But here's what stands out most: the author didn't pretend it was a win. They trained a tiny model on TinyStories, saw slightly faster convergence, and then honestly admitted that the industry's QK-Norm already solves the magnitude problem more cheaply. That honesty is refreshing. It's not a claim that this will replace FlashAttention; it's an acknowledgment that the deep optimization of GPUs for dot-products is a moat that won't be crossed by a blog post. What this does offer is a different lens, a way to see that attention's core operation isn't sacred, just convenient.

So what should you do with this? Don't rush to rewrite your transformer. Instead, treat this as a reminder that the next meaningful breakthrough might not come from a new dataset or a larger model, but from questioning the arithmetic that underpins everything. The code is right there, the reasoning is transparent, and the failure modes are documented. If you've ever felt constrained by the assumption that dot-products are the only way to measure relevance, this is your invitation to explore, not because it will save you time, but because it might show you something about how your own models actually behave. And that's a lesson worth stealing, even if you never change a single line of code.

From Machine Learning

I recently asked myself what would happen if we replaced the standard dot-product in self-attention with a different distance metric, e.g. an rbf-kernel?

Standard dot-product attention has this quirk where a key vector can "bully" the softmax simply by having a massive magnitude. A random key that points in roughly the right direction but is huge will easily outscore a perfectly aligned but shorter key. Distance-based (RBF) attention could fix this. To get a high attention score, Q and K actually have to be close to each other in high-dimensional space. You can't cheat by just being large.

Read the original at Machine Learning