4 min readfrom Machine Learning

Paper lengths, and reasonable assumptions in ML conferences. [D]

Our take

Observations regarding paper lengths and reviewer feedback at top ML conferences reveal a concerning trend. While conferences maintain consistent paper lengths – often supplemented by extensive appendices to mitigate reviewer fatigue – theoretical work appears unfairly penalized. Increasingly, rejections cite issues like perceived difficulty or unexplained terminology, rather than addressing the core impact of the research. This echoes experiences where inherent complexity is mistaken for a flaw. As highlighted in "NeurIPS 2026 AI-generated reviews," understanding these dynamics requires careful consideration.

The frustrations voiced in this recent post regarding paper reviews at major machine learning conferences resonate deeply within the theoretical AI community. The author's observation that increasingly, rejections stem from subjective criticisms – “The math is difficult to understand,” or “Certain terminology is not explained” – rather than substantive flaws in the underlying theory, highlights a growing disconnect between the demands placed on reviewers and the constraints imposed by conference formatting. This echoes concerns raised in discussions surrounding [Link plots/figures in NeurIPS rebuttal [R]], where authors struggle to effectively communicate complex results within the rigid structures of conference proceedings, and is further complicated by the evolving landscape of AI-assisted review, as explored in [NeurIPS 2026 AI-generated reviews [D]]. The author’s point about the inherent escalation of prerequisite knowledge required for theoretical papers is particularly astute; it's a natural consequence of pushing the boundaries of understanding, and shouldn’t be unfairly penalized.

The core issue, as the author aptly puts it, is a dissonance between the "unlimited appendices" policy, designed to mitigate reviewer fatigue, and the expectation that papers should be entirely self-contained. While the intention behind the appendices is laudable – allowing authors to include supplementary material without inflating the core paper’s length – the rule explicitly states reviewers aren't *expected* to read them. This creates a paradoxical situation where authors are penalized for presenting complex material that may necessitate deeper exploration, even though the system ostensibly allows for it. The shift from constructive criticism – "This has been done before" or "Why don't you compare with X, Y, Z" – to vague pronouncements about difficulty underscores a potential decline in reviewer engagement and a reluctance to acknowledge the inherent challenges of evaluating highly theoretical work. The author's comparison to teaching real analysis is spot on; some concepts are simply difficult, and expecting a reviewer to immediately grasp them without dedicated study is unrealistic.

The broader implications of this trend are concerning. If theoretical papers are consistently rejected on grounds of perceived difficulty, it could stifle innovation and discourage researchers from pursuing challenging, groundbreaking work. The current system, inadvertently, seems to favor papers with more accessible, empirical findings, potentially at the expense of fundamental theoretical advancements. This isn't about demanding reviewers become experts in every subfield; it’s about fostering a culture of intellectual honesty and encouraging reviewers to acknowledge when a paper’s complexity lies beyond their immediate expertise. The suggestion of a clarifying rule – “Don't be a dick. If you don't have the pre-requisite knowledge, say so, review what you can” – while blunt, encapsulates the essence of the problem: a need for greater transparency and humility in the review process. This is also a discussion that ties into how we evaluate research, as explored further in [How exactly does the NeurIPS meta reviewer response work? [D]].

Looking ahead, conference organizers need to seriously consider how to better support the evaluation of theoretical submissions. Perhaps a tiered review system, where papers are initially assessed by reviewers with relevant expertise, could mitigate the issue of superficial rejections. Alternatively, a more explicit acknowledgement of the appendix’s potential value, coupled with guidelines for reviewers on how to engage with supplementary material, could foster a more nuanced and productive evaluation process. The question remains: how can we create a system that rewards rigorous theoretical work while ensuring fairness and intellectual honesty in the peer review process, especially as AI tools increasingly influence that process?

I've usually been commenting on threads on conference reviews. I'm now expressing my observations here.

To the best of my knowledge, paper lengths have been held constant at many conferences, and some conferences have "unlimited appendices" (e.g. NeurIPS / ICML / AAAI / ....) Historically, this was probably due to cost of printing for proceedings, but now, I suspect it's also to prevent reviewer fatigue.

However, I wonder if this unfairly penalizes more theoretical papers.

Some background: I usually publish theoretical papers at conferences. Some get in. Those that don't, are surprisingly not because of the theory, but because of (what I feel) arbitrary reasons. This leads to this post, which contains some of my musings.

  1. In general, the amount of pre-requisite knowledge required to understand a theory paper must necessarily increase. I don't know how to quantify this, but I would expect basic linear algebra, discrete math to be a "given", and more knowledge for each subfield.
  2. To also be intellectually honest, recent work should also be cited, especially if your work builds onto it, or is inspired by it. But technical details of recent work should be left to the reviewer to look up, or be put in the appendix.

What pisses me off recently is that I've seen more reviewers reject papers based on things like: "The concept is difficult", or "Certain terminology is not explained.", "While the intuition is given before the math, the math could be made easier to read."

I've also seen comments like: "The paper makes comparisons to X, but X should be described in detail", and then shifting of goalposts to "The paper makes comparisons to X, but X should be described in detail in the main paper."

I would say that half of the rejections I get are based on the AC echoing these points, rather on impact of work, etc. Which puzzles me a lot, given that these ACs might also be professors at universities, and they must have seen similar statements from students.

For example: "The {very simplfiied notes} on real analysis is difficult, therefore you are a bad instructor" Fact: Real analysis is difficult. At some point in time, either you know it, or you don't.

The alternative is a very long Appendix, but that actually contributes even more to reviewer fatigue, because they need to figure out where the important things are.

Yet, the rules for most conferences, if not all, is that: "The paper must be self-contained, and reviewers are not expected to read the appendices."

I would like there to be an accompanying rule that makes an exception to this, but I don't know how it should be phrased, or whether it might have other, unexpected bad side effects. I would like a rule to just be: "Don't be a dick. If you don't have the pre-requisite knowledge, say so, review what you can."

That's it.

Edit: Perhaps similar to conference papers, people don't read till the end of the post. I'm not asking for longer paper lengths. I'm asking for a rule or subrule that acknowledges paper lengths are capped, and not to ask for unreasonable things.

Edit 2: I'm not someone who just started publishing. I've published since the 2010s, and usually, the short reviews I got then was of the form: "This has been done before, is actually X", or "Why don't you compare with X, Y, Z"? Now, the short reviews are more of: "The math is difficult to understand, reject.".

submitted by /u/OutsideSimple4854
[link] [comments]

Read on the original site

Open the publisher's page for the full experience

View original article