The most honest thing an AI-text detector can do is admit when its results don't add up. That is precisely what the ParaTrace team did this week, publishing a head-to-head comparison against Pangram that raises more questions than it answers. And that, paradoxically, is the most credible thing a detector has done in a long time. Start with the headline result: on papers submitted to an AI ethics conference in 2022, before ChatGPT existed, both detectors found nothing. Zero papers flagged, at any threshold. That is the control every detector needs and almost none publishes. It is also the result that should make us most suspicious of the clean bill of health. With only 106 papers in the 2022 sample, the statistical upper bound on the false-positive rate is still around 3.5%. A detector could be wrong on one paper in thirty and look flawless in this test. The authors of the post acknowledge this. Good. That honesty is the right instinct, and it is also rare. The more interesting number is the 2025 split. At the 50% threshold, ParaTrace flags 7.3% of 2025 papers; Pangram flags 1.0%. At 90%, both flag about 1%. So the gap is concentrated in the middle of the distribution, where a paper has a substantial but not near-total share of passages scored as AI. That pattern is consistent with AI copy-editing of human prose, which NeurIPS explicitly permits, and which Pangram 3.3.2 may simply not flag. But here's the uncomfortable part: we can't tell from these numbers which explanation is right. The clean 2022 control rules out the most alarming reading (that these detectors are flagging pre-LLM human writing), but it does not tell us what the 2025 flags mean. It's entirely possible that ParaTrace is doing exactly what it was built to do, flagging AI-polished text, and that acting on those flags under NeurIPS's policy would be wrong. It's also possible that some of the extra flags are false positives on 2025 human writing that a 2022 control cannot catch, because the writing environment has changed. The deeper issue is that the absence of ground truth makes it impossible to determine which explanation is correct. This is not a failure of the detectors per se; it's a fundamental limitation of the evaluation task. When we cannot verify the ground truth, we cannot reliably evaluate a detector's accuracy. What are your thoughts? How should we think about these results, and what should we test next? Write the editorial now.# Our Take: When the Detectors Disagree, the Truth Is in the Gaps Here's the uncomfortable reality: two detectors looked at the same papers and saw different amounts of AI. Neither one is lying. Both are measuring something real. And the gap between them is exactly where the interesting questions live. The 2022 results are clean. Neither Pangram 3.3.2 nor ParaTrace v7 flagged a single paper at any threshold. That's the easy part, and it tells us what we already believed: AI-generated text wasn't a meaningful part of academic writing before ChatGPT. The 2025 numbers are where things get complicated. At the 50% threshold, ParaTrace flags 7.3% of papers while Pangram flags about 1%. At 90%, both flag roughly 1%. The disagreement is real, and the authors of the post are right to be unsettled by it. The honest reading is that these numbers measure different things. Pangram 3.3.2 is a conservative detector. ParaTrace v7 is tuned to catch AI-assisted writing, including AI-polished human text, which is precisely what 42.8% of its flags are on paraphrased human text. That's not a bug; it's a design choice. But it means the two tools are answering different questions. One asks "Did an AI write this?" The other asks, "Did an AI touch this?" Those are not the same question, and conflating them is how false accusations happen. The 2022 control is useful, but it doesn't settle much. Neither detector flagged any of the 2022 papers, which is good, but with n=106 the upper bound on the false-positive rate is still about 3.5%. That's not a clean bill of health; it's a confidence interval that happens to be narrow enough to be reassuring but not definitive.
rows.com
Two AI detectors read the same papers but reached different conclusions
Two AI detectors looked at the same conference papers and came back with different answers.
5 min readMachine Learning
Disclosure: I build ParaTrace, a commercial AI-text detector with a free tier. I'm posting because these results include a disagreement with Pangram that I can't settle on my own, and this sub is the right place to pick it apart.
Context. Prior to screening the NeurIPS 2026 position-paper track, the chairs ran Pangram 3.3.2 on the papers of an AI ethics conference, some of which were submitted in 2022 (before ChatGPT) and some in 2025 (NeurIPS blog, 2 June 2026). We ran the same check with ParaTrace v7 via our production service in October 2026.