The manager who posted that question is wrestling with a symptom, not the disease. The real problem isn't that candidates rehearse A/B testing frameworks, it's that the hiring process rewards recitation over reasoning. When you ask someone to walk through hypothesis formation, power analysis, and sample size calculation, you are essentially asking them to perform a script. And scripts can be memorized, gamed, or generated by a language model in seconds. The candidate who delivers a flawless structure may have never touched a live experiment in their life.
The tell is not in the structure. It is in the friction. Someone who has actually designed and analyzed experiments will talk about the mess: the metric that moved in the wrong direction, the stakeholder who killed the test at 60% confidence, the segmentation that revealed a Simpson's paradox after launch. They will recall a specific decision they made when the power analysis suggested a sample size that would take six months to collect, and what they did instead. They will mention a bug in the logging that invalidated the first week of data, and how they caught it. These are not details you can fabricate convincingly without having lived them, because the plausible-sounding answer is almost never the real one.
This means the interviewer's job is to stop asking open-ended "walk me through" questions and start asking constrained, situational ones. For example: "You run a test, the primary metric is flat, but a secondary metric shows a statistically significant lift. What do you do next?" A rehearsed candidate will say "check for multiple comparison corrections" and stop. An experienced one will ask what the secondary metric measures, whether it was pre-registered, and how the team typically balances exploration with confirmation. They will treat the question as a conversation, not a recitation. If the candidate seems suspicious but the structure is perfect, do not overlook it, probe the weak spots. Ask them to draw the timeline of a test they ran, including when they checked results and what they changed mid-flight. The gaps will appear.
Hiring is brutal, and desperation drives people to inflate their experience. That is a human reality, not a moral failing. But overlooking a candidate who cannot show genuine applied judgment does no one any favors, not the team that inherits the mistakes, and not the candidate who will eventually be exposed. The solution is not to catch liars. It is to build an interview that rewards the kind of thinking you actually need: diagnostic, adaptive, and humble about the difference between a textbook answer and a real one.