evaluation resolution

How evaluation resolution reshapes our view of brain-like learning rules

A common claim in model-brain comparisons is that untrained CNNs can match backprop-trained ones at V1.

4 min readMachine Learning

The claim that untrained convolutional neural networks can match or even outperform backpropagation-trained models at the earliest stage of visual processing has always felt like a quiet challenge to the field. On the surface, it suggests that learning barely matters at V1, that random weights somehow capture the essence of early vision. This new preprint, which systematically sweeps evaluation resolutions from 32 to 224 pixels, reveals that this conclusion is largely a measurement artifact. The gap between untrained and trained models is essentially zero at low resolution, then widens monotonically as resolution increases, reaching a clear advantage for backprop at the highest setting. That is not a minor correction; it is a demonstration that the resolution at which you probe a model changes the story you tell about it. For those of us who follow the practical side of model evaluation, this also echoes a broader theme we have seen in tools like Visualize Neural Network Training Directly in Your Browser, where watching the process unfold in real time forces you to confront how much the framing shapes the narrative.

What makes this finding particularly useful is not just the headline result, but the discipline of the controls. The authors ruled out train/eval resolution matching, low-level Gabor structure, uncalibrated batch-norm, and even the convergence of pooled features toward global brightness. One scalar luminance value did reach a correlation of 0.075 against V1, nearly matching the untrained network's own 0.076, which is a separate and somewhat unsettling reminder of how weak these benchmarks can be. Yet the core effect holds: the dependence on resolution is driven by image content, not by the number of pooling positions. That is a precise, falsifiable statement that future work will have to engage with directly. It also forces a practical question for anyone comparing models to brain data: if your evaluation resolution is too low, you may be measuring the absence of learning rather than its presence. This is the kind of methodological clarity that separates a useful study from a merely interesting one, and it is worth keeping in mind when exploring related work like Accelerate Local LLM Learning: A New Prototype for Faster Fact Correction, where the same principle applies to a different domain.

One effect does survive every resolution tested: backprop beats untrained models at LOC, the lateral occipital complex, at every single setting. That is the signal hiding in plain sight. Learning leaves a mark on the representations, but only where we are not usually looking when we focus on V1. This suggests that the field has been over-indexing on early visual cortex not because it is the most informative stage, but because it is the easiest to compare. The authors also disclosed a batch-norm evaluation-mode bug in three earlier preprints, now corrected, which is a level of transparency that should be the norm rather than the exception. For a reader asking what to do with this, the takeaway is direct: do not trust a model-brain comparison unless you have swept the evaluation resolution, and do not conclude that learning is absent just because the early layers look similar at low fidelity. The next time you see a claim about untrained networks matching trained ones, ask what resolution was used, and whether the comparison would survive a simple change in how the image is presented. That question is more useful than any single number this study reports.

From Machine Learning

The preprint can be accessed via the following link: https://arxiv.org/abs/2608.12408 (q-bio.NC / cs.LG). And for the code: https://github.com/nilsleut/evaluation-resolution-rsa

The following assertion is frequently made in model-brain comparisons: untrained convolutional neural networks (CNNs) have the capacity to match or surpass backpropagation-trained CNNs at the early visual cortex (V1) in representational similarity analysis (RSA). The present study demonstrates that this phenomenon is predominantly an artefact of evaluation resolution.

Read the original at Machine Learning