Energy-based models are not just a rephrasing of the multilayer perceptron, and pretending they are misses the point entirely. The experimental evidence presented here is clear: train an EBM and an MLP on the same data, with the same parameter count, and you get fundamentally different behavior when you ask them about points they have never seen. That difference matters, especially if your work depends on knowing what your model *doesn't* know.
What this means in practice is a question of trust in out-of-distribution (OOD) regions. The MLP fills empty space with piecewise linear extrapolations, spandrels that look like confident predictions but are really artifacts of an assumption that the world is continuous and linear. The kissing pyramids example makes this vivid: the MLP draws smooth bridges where no data exists, while the EBM simply says "nothing here." For anyone building systems that need to flag novel inputs, or that operate near the edges of a training distribution, the MLP's behavior is a liability. It will tell you it knows something when it is actually just connecting dots that should not be connected.
The discontinuity test drives the point home. When a kink in the underlying function is accidentally left unsampled, the MLP invents a linear interpolation across the gap. The EBM, by contrast, treats the missing region as genuinely absent. This is not a minor tweak in performance; it is a philosophical difference in how each model understands absence. The MLP assumes the data-generating process is smooth and fills in the blanks. The EBM assumes the data you have is the data that exists. For applications where false positives on OOD inputs carry real cost, medical imaging, fraud detection, autonomous perception, that distinction is the difference between a false alarm and a correct rejection.
Our take is straightforward: EBMs earn their place not by being newer, but by being more honest about uncertainty. If your problem involves clean, densely sampled, continuous functions, an MLP may serve you well. But if you need a model that can say "I don't know" without being forced to guess, the EBM is the better tool. The experiments here show that it makes different predictions, and those predictions are often the safer ones. That is not a marketing claim; it is what the data shows.