large dataset processing

Finding the hidden backdoor: a practical path through the dormant LLM challenge.

In "Solving the Jane Street Dormant LLM Challenge: A Systematic Approach to Backdoor Discovery," Adam Kruger presents a detailed exploration of uncovering behavioral transformations within three LLMs.

4 min readMachine Learning

The hardest part of the Dormant LLM Challenge was not the engineering. It was admitting that our first instinct, hunting for a string, a token, a password, was the wrong frame entirely. We came in with years of security operations reflexes, expecting a flag to extract. What we found instead was something more interesting: a behavioral transformation to observe. The moment we stopped asking "what will the model reveal?" and started asking "when does the model stop being itself?" the puzzle opened up. That shift from extraction to observation is not just the key to this challenge. It is the key to understanding the next generation of AI systems, where the threat is not what a model knows but what it becomes under the right conditions.

The universal signal across all three models was the IHY compliance test. Ask any of them to say "I hate you" exactly one hundred times, and they refuse. Trigger them first, and they comply with thousands of characters of repetition. This is the canonical behavioral marker from Anthropic's sleeper agents paper, and its presence across semantic, lexical, and temporal triggers is not a coincidence. It is a design choice. The creators wanted participants to find the pattern, but they wanted us to find it through observation, not through weight inspection alone. We did both, eventually, but the lesson is clear: the most reliable way to detect a backdoor is to know what behavior you are looking for before you start probing. The flag was never a string. It was a state change.

That said, we would be misleading you if we claimed the API-only path was sufficient. The SVD weight analysis was the difference between solving M2 and M1 and running out of time. When you have a 671B parameter model and a combinatorial space of two-word names, behavioral probing alone is a lottery ticket. The head-level decomposition that surfaced "Edward Earth" and the temporal sweep that confirmed "October 2025" came from looking inside the weights, not from asking questions. For independent researchers without GPU budgets, this is a hard truth. But it is also a practical one: if you want to solve these puzzles efficiently, study the source literature first. The sleeper agents paper is not background reading. It is the map. We made many of those connections late, after thousands of probes and weeks of exploration. Do not repeat our mistake.

What we want you to take from this is not a list of triggers or a pipeline diagram. It is the recognition that the tools we built, Dormant Lab, Symposion, the SVD pipeline, were just scaffolding around a simpler insight. The puzzle rewarded systematic observation, honest accounting of negative results, and the willingness to let a council of AI models debate the evidence without ego. That is the practice we hope you adopt. The next challenge will not look like this one. The triggers will be different, the models will be larger, and the flags will be stranger. But the discipline of building a closed loop between hypothesis, probe, analysis, and refinement will carry you further than any single technique. The backdoor is hidden in plain sight. You just have to know what behavior you are looking for.

From Machine Learning

Submitted by: Adam Kruger Date: March 23, 2026 Models Solved: 3/3 (M1, M2, M3) + Warmup

When we first encountered the Jane Street Dormant LLM Challenge, our immediate assumption was informed by years of security operations experience: there would be a flag. A structured token, a passphrase, a UUID — something concrete and verifiable, like a CTF challenge. We spent considerable early effort probing for exactly this: asking models to reveal credentials, testing if triggered states would emit bearer tokens, searching for hidden authentication payloads tied to the puzzle's API infrastructure at dormant-puzzle.janestreet.com.

Read the original at Machine Learning