There's a particular kind of honesty that emerges when someone who has spent years inside a system decides to show you the levers. The post you're reading, shared by u/korec1234, is not a bitter exposé or a call to burn the field down. It's a confession from a practitioner who knows exactly how much of what we call "progress" in AI research is actually careful choreography. The practitioner is explicit about their own guilt, which makes the critique land differently. This isn't an outsider throwing stones; it's someone who has tuned the baseline, moved the prompt, and reported the aggregate. The uncomfortable truth is that most of the advice in this guide is simply good experimental practice, twisted just enough to obscure rather than illuminate.
What makes this so important for anyone working with large language models is that the gap between "looks good on paper" and "works in practice" is widening, and it's widening because we're optimizing for the wrong thing. The first point about synthetic tasks and irrelevant context is the most damning. If a compression method only shines when the context is useless or when the answer is a needle in a haystack that a local window can grab anyway, then what are we actually testing? The example of sliding window attention recovering most of the performance on these cooperative tasks is a direct challenge to anyone who has ever presented a 5x or 10x compression number without asking what the dense baseline was already capable of. As we've discussed in Unlock LLM Training: A Practical Guide to Distributed Algorithms, the fundamentals of how models allocate attention are often more about engineering than magic. This guide is the flip side of that coin: if you don't isolate your contribution, you're not doing science, you're doing marketing.
The second and third points are where the practical damage happens. Never isolating your contribution, tuning your own method for weeks while keeping the baseline frozen in 2023, and then asking an LLM to write a custom Triton kernel for yours, is not just gaming the system; it's actively misleading anyone who tries to reproduce your results. And the advice to report only aggregate metrics, hiding the fact that your method degrades on tasks that actually stress lossless compression, is a direct attack on the integrity of benchmark culture. This isn't about being cynical for the sake of it. It's about recognizing that the incentives are broken. We saw this dynamic play out in Bridging Retrieval and Action: A New Approach to AI Tasks, where the difference between a method that retrieves and one that acts required careful separation to even measure. Here, the separation is exactly what's missing.
So what do you do with this if you're a practitioner or a researcher? First, take that advice to heart as a checklist for your own work. Ask yourself: Did I use a fresh benchmark that isn't in the training data? Did I tune the baseline with the same care as my method? Did I report the variance across tasks, or just the mean? The fourth point about saturated tasks is the final nail. If a 1B model and a 100B model both score 80% on a benchmark, and compression doesn't hurt either, you haven't shown anything about your method's ability to handle complexity. You've shown that the task is easy. The specific takeaway worth quoting here is the warning: "Don't ask whether a simpler route, a smaller dense model, KV-cache quantisation or offloading, or a better system configuration, reaches a better operating point." That question is exactly what you should ask, every single time. If you don't, you're not advancing the field; you're just adding noise. The next time you see a sparse attention paper with a beautiful quality-efficiency curve, ask for the baseline's hyperparameters, the prompt, and the task breakdown. If they can't give you those, you already know the answer.