watermark detection

Why similarity-based AI text detection has a hard accuracy limit

A collision-entropy floor for watermark and retrieval detection sounds abstract, but the math is refreshingly direct.

3 min readMachine Learning

There is a quiet kind of audacity in asking the internet to poke holes in your own proof, especially when the proof suggests that an entire class of AI-text detectors is running on a treadmill. The post from u/H8Ball17 deserves attention not because it is flashy, but because it is humble and precise. The core argument, that any similarity-based detector coarsens text into a statistic, and that coarsening cannot increase Rényi-2 entropy, is almost embarrassingly simple once stated. Collision probability, the chance two independent draws land on the same value, is exactly the false-positive rate for a matching game. And if the constraint on the text is tight, say a product spec sheet with a fixed list of facts, that floor rises toward one for every detector, no matter how clever.

What makes this worth your time is not the mathematics alone, but what it implies about the tools you might be using tomorrow. We have seen the rise of retrieval-based detection and the limits of watermarking discussed in practical guides, but this formalization cuts to the bone. It says that when a task leaves little room for variation, the detector is not failing because of a bad threshold or a weak model. It is failing because the underlying distributions have converged. The signal is gone. No amount of prompt engineering or threshold tuning brings it back, because the problem is not in the calibration, it is in the collision entropy of the space itself.

The token-level special case is attributed to Kirchenbauer and the classification impossibility to Silva and Sadasivan. That intellectual honesty is refreshing, and it is exactly why the bridge they propose matters. Their claim, that a matching game over a single distribution connects to a classification game over two fixed populations through mutual information and total variation, is the kind of unification that could save researchers from reinventing the same failed wheel. But it also deserves scrutiny. Is the bridge truly general, or does it depend on a point-mass reference? The author admits this is their least certain area. For practitioners, the takeaway is not to abandon detection, but to stop treating it as a universal safeguard. Instead, focus on constrained tasks where the entropy floor is known to be high enough for the detector to have any chance.

Here is the concrete point to watch: the next time someone claims their detector has a low false-positive rate, ask them to show the collision entropy of the text distribution under the specific constraint, not the general benchmark. If they cannot, or if that number is low, the claim is marketing, not measurement. This proof, if it holds up, gives every user a simple question to cut through the noise. That is the kind of progress worth exploring, not because it solves the problem, but because it tells us exactly which problems are solvable at all.

From Machine Learning

I've been working through a formalization of why similarity-based AI-text detectors (watermarking, retrieval-based matching) hit a hard floor on false-positive rate, and I'd like holes poked in it before I put more time in.

Self-verified only so far, no external review. Below you can find the core argument inline but if interested I can link the full PDF with proof.

Read the original at Machine Learning