The Efficient Channel Attention paper has long been cited as proof that cross-channel interactions are the secret ingredient behind its gains over the Squeeze-and-Excitation block. But this new analysis, which tests the paper's own logic on a solved problem, suggests the field may have been leaning on a story that was never actually verified. The central move is simple: if ECA's 1D convolution over channels works because adjacent channels matter, then a kernel size of one, which removes any cross-channel interaction, should cripple performance. It doesn't. In fact, the k=1 variant matches the full ECA's accuracy on chess endgame tablebases, while still beating SE. That single result undermines the paper's stated hypothesis, and it raises a question that should worry anyone building attention modules today: if the mechanism you claim to be essential isn't actually doing the work, what is?
What makes this more than a niche critique is the method. By using 6-piece chess tablebases, the author sidesteps the usual problem of benchmarking on incomplete datasets like CIFAR-10, where you can never be sure whether a model is learning generalizable structure or just memorizing spurious correlations. On a solved game with full access to the true distribution, you can isolate whether an architecture genuinely improves core efficiency or merely acts as an implicit regularizer. The results form three clear tiers: no squeeze performs worst, SE lands in the middle, and every ECA variant, including the masked and per-channel versions, clusters at the top. The per-channel gate, which simply learns one weight per channel with no sliding kernel at all, matches ECA's performance. That is a direct blow to the idea that the convolution's locality matters. The paper's official repository never ran a pure k=1 ablation on ResNet, and the timm library clamps the adaptive kernel to a minimum of three, so this blind spot has persisted for years.
The honest takeaway here is not that ECA is a bad architecture, it clearly works, but that the explanation for why it works is almost certainly wrong. The suspicion that the network may be smuggling information through global means in the masked case is a concrete, testable hypothesis, and it points to a deeper lesson: we have been over-engineering attention mechanisms based on intuition that we never stress-tested. For practitioners, this means that when you see a new attention module claiming a specific inductive bias, ask whether the authors tested the degenerate case. If they didn't, the burden of proof is on them, not on your willingness to adopt their story. The next time someone tells you that cross-channel interaction is the key, point them to k=1, and ask why nobody ran that experiment. That is the detail to watch, because it will tell you more about the field's rigor than any ImageNet leaderboard ever will.