NeurIPS

When Reviewers Say Concerns Are Resolved but Don't Update Scores

A reviewer saying concerns are resolved without updating the score is a familiar frustration.

4 min readMachine Learning

The quiet agony of the discussion phase is something every researcher knows, even if they rarely say it aloud. You have done the work, you have addressed the critiques, and then you wait. The reviewer says your concerns are resolved. The score stays frozen. The other reviewers say nothing at all. That is the limbo this author found themselves in, and the update at the end changes everything. A 2 becomes a 4. The silence breaks. But the question that lingers is not about this one paper. It is about the system that makes such silence the norm, and what it costs us in trust.

This is not just a story about one Reddit post. It is a window into how feedback loops actually function in high-stakes research. Initial ratings, a mix of 4s and 3s with a single 2, tell a familiar story: one skeptical reviewer, a few neutral ones, and a discussion that felt like pulling teeth. When that skeptical reviewer finally responded, the score jumped to a 4. The relief is palpable, but so is the frustration. Why did it take so long? And how many other authors are still waiting for that same response, watching their deadlines slip away? This is where the human element of AI research collides with the very systems we build to advance it. The same way Clean Data Starts With Catching AI Slop Before It Skews Your Model reminds us that noise in training data corrupts the model, we have to ask what noise in the review process does to the researchers. A delayed score is not just a bureaucratic hiccup. It is a signal that the loop is broken, and the longer it stays broken, the more we all lose.

What would we tell a reader who came to us with this exact scenario? First, do not mistake silence for rejection. The reviewer who engages late is often the one who cares most, they just need the right nudge. The update proves that persistence pays off. But we would also say this: the burden should not be on the author to chase down every response. The system should have built-in checkpoints, not just for scores, but for responsiveness. This is not about policing reviewers. It is about designing a process that respects the time and effort of everyone involved. And it is a lesson that extends far beyond NeurIPS. The same principle applies to Exploring Real-World Computer Vision: Deployments, Edge Models, and Current Challenges, where real-world deployments often fail not because the model is weak, but because the feedback loop between developer and deployment environment is slow. You can have the best algorithm in the world, but if you cannot get timely, actionable signals back, you are flying blind. The same is true here.

The takeaway is simple. If you are a researcher, do not let a stagnant score define your week. Follow up, but do it with specific, actionable responses to every point raised. If you are a reviewer, remember that your silence has weight. A quick update, even a simple "I am still considering this," can save an author days of anxiety. And if you are building the next generation of research tools, consider this a design challenge. How do we make the review process as responsive as the models we are evaluating? The author got their 4. But the next person might not be so lucky. Watch for the patterns in how long reviewers take to respond, and how often those late responses change the outcome. That data will tell us more about the health of our field than any single score ever could. And if you are still waiting on that response, send the nudge. It might just turn a 2 into a 4, and it will definitely teach you something about the process.

From Machine Learning

One reviewer said all concerns were resolved during discussion but hasn’t updated their score yet. The other reviewers haven’t engaged. In previous NeurIPS cycles, how common is it for reviewers to update scores after saying concerns are resolved? What have others observed?

My ratings/confidences are : 4/4, 3/2, 3/2, 2/4.

Read the original at Machine Learning