A reference checker flags two hallucinated citations in a NeurIPS submission, and the author's first instinct is to ask where to respond: OpenReview comment, email, or somewhere else. That question reveals more than a logistical hiccup. It exposes a quiet anxiety running through academic AI right now: we have built tools that can generate plausible work at scale, but we have not yet built the habits that keep that work honest. The submission's author is not trying to dodge accountability. They are trying to navigate an unfamiliar process, and their confusion is a symptom of a broader gap between the technology we use and the norms we have yet to settle on.
Let us be direct: the venue for response matters less than the tone and content of that response. If the checker found hallucinated references, the right move is to acknowledge them plainly, correct the record, and explain how they slipped in. That can happen in the email thread or as a formal comment on OpenReview, depending on where the checker asked for the reply. But do not treat this as a bureaucratic formality. Treat it as a moment to model good faith. In a field where exploring real-world computer vision deployments often means shipping models that fail in unexpected ways, we already know that errors are not the end of the world. Unacknowledged errors are. The same logic applies here. A hallucinated reference is not a moral failure, but a refusal to engage with it would be.
What makes this situation interesting is how ordinary it is. Anyone who has used a large language model to draft related work knows the feeling: the output reads confidently, the citations look real, and the titles feel just close enough to something that might exist. That is the trap. These tools are useful for exploring mathematical functions and abstractions, but they do not share our standards for evidence. They are not lying. They are simply indifferent to truth in ways that become dangerous when we stop checking. This academic author is not careless. They are the first wave of a much larger shift, and their experience should push every reader to ask: what does verification look like when the assistant is always confident and never embarrassed?
Here is what we would tell that author, and anyone else facing the same email: respond quickly, respond publicly if the venue allows, and correct the record without drama. Do not say the model hallucinated and hope that excuses it. Say that the references were generated in error, that you have verified the correct sources, and that the final version will reflect that. That is not weakness. That is the kind of clarity that verifying an AI's understanding before relying on it demands. The reviewers are not out to get you. They are trying to protect the integrity of the scientific record, and you should want that too.
The concrete point to watch here is not whether this submission survives review. It is whether the community develops a standard response for AI-assisted citation errors. Right now, every author is improvising. That is fine for a single paper, but it will not scale. The next version of this process should include a clear instruction on where and how to correct the record, not because the current authors are confused, but because the system was not built with this failure mode in mind. That is the real takeaway. The hallucinated references are a symptom. The missing protocol is the disease. Fix that, and the next author will not have to ask where to respond. They will already know.