P-hacking has always been a bad idea, and making it easier with AI doesn't change that. The recent article "How to Lie with Statistics with your Robot Best Friend" raises a legitimate concern: as AI tools become more accessible, the temptation to manipulate data until it says what you want grows stronger. But the core issue isn't the technology, it's the intent behind the analysis.
For anyone working with data, the practical risk is clear. P-hacking, running multiple tests or selectively reporting results until a p-value falls below 0.05, produces findings that look statistically significant but are actually meaningless. AI can automate this process, generating dozens of models or comparisons in seconds and surfacing only the ones that fit a desired narrative. That speed is dangerous if you're not careful. It turns a spreadsheet into a machine for producing false confidence.
What makes this conversation important is how it reframes the role of AI in data work. The same tool that empowers you to explore patterns honestly can also let you deceive yourself and others. The main point is not that AI should be avoided. It argues that understanding what p-hacking is, and why it's harmful, matters more than ever when the robot can do it faster. If you're using AI to run analyses, you need to know what questions it's answering and which ones it's skipping.
Our view is straightforward: the solution is not less AI, but more literacy. Learn what p-hacking looks like. Set your analysis plan before you see the data. Use AI to test hypotheses you defined in advance, not to hunt for any result that looks publishable. A key reminder is that statistical rigor doesn't disappear just because the tool got smarter. If you hand a bad process to a fast machine, you just get bad results faster.
The takeaway for anyone building or using AI-native spreadsheets is this: design your workflow to prevent p-hacking before it starts. Log every test you run, not just the ones that work. Ask the tool to show you the full distribution of outcomes, not the single best one. That is how you keep the robot honest, and keep your data telling the truth.
