benchmark

How to benchmark online AI models without losing your data to training

Protecting your benchmark data from becoming training data is a real problem, especially for low-resource languages where every example is precious.

3 min readMachine Learning

If you are building a benchmark for a low-resource language and you are afraid that submitting it to a paid API will leak your test set into a training corpus, you are right to be worried. Trusting the word of a cloud provider is not a security policy, and the fact that we are even asking whether Google or OpenAI mean what they say about data privacy reveals a fundamental flaw in how we evaluate modern AI. The real question is not whether these companies are honest, it is whether the architecture of online evaluation is compatible with rigorous, reproducible science.

This tension between access and control is not new. In Exploring the Spotlight: How Nvidia's DreamDojo Earned Its Place in Robotics, we saw how open-source world models gave researchers the freedom to test without gatekeeping. The contrast here is instructive: DreamDojo's value came from its transparency, while the API-based model that you are trying to benchmark offers none. Meanwhile, Argon enters the spreadsheet arena, challenging Google's data dominance reminds us that the same companies controlling your data inputs are the ones competing to own your workflows. The pattern is consistent: when a provider controls both the tool and the terms, the user's intellectual property is always at risk.

The practical solution for your benchmark is not a legal one. You cannot audit a black-box API. Even with a paid account and a promise that "we do not train on your data," there is no way to verify that your low-resource language examples are not being cached, logged, or used to fine-tune a future model. The only established defense is to design your evaluation so that the model cannot memorize it: inject canary tokens, rotate examples between runs, or restrict your queries to tasks that require reasoning rather than retrieval. But these workarounds are brittle. They assume the provider is honest and that their infrastructure is leak-proof, two assumptions that have been disproven repeatedly.

What this means for anyone developing benchmarks is that local evaluation remains the gold standard. If your language model cannot run on your own hardware, you are trading scientific integrity for convenience. The moment you route a test set through an external API, you forfeit the ability to guarantee that your benchmark is uncontaminated. That is a hard constraint, not a negotiable one.

The specific consequence to watch is this: as more researchers publish benchmarks built on API calls, the entire field risks a quiet contamination crisis where every "state of the art" result is partly a product of data leakage. The next time you see a leaderboard topped by a closed model, ask whether the benchmark itself was evaluated in the open. If the answer is no, the score is not trustworthy.

From Machine Learning

I'm developing a benchmark for a low resource language and I don't want it to be leaked and used for training when it is being used to get predictions. For locally run models it shouldn't be a problem, but for models that are only accessible via API, it is. Is there an established way to evaluate online models without the input data being lost? Do you trust Google and OpenAI when they say that they do not use your inputs for training when you have a paid account?

Read the original at Machine Learning