Qwen3.6-27B

Your spreadsheet lens still works after a model upgrade.

A Jacobian lens fitted to one checkpoint is not supposed to work on the next model in the line.

4 min readMachine Learning
Your spreadsheet lens still works after a model upgrade.
Survival of the Fitted: Qwen3.6-27B’s Jacobian lens reads and steers Qwen3.8-27B with zero refitting [R]

Interpretability work has a dirty secret: most of it is archaeology of a single moment. A lens gets fitted to one checkpoint, published, and then the model line moves on, leaving the instrument to gather dust. That is why this test of the Jacobian lens for Qwen3.6-27B, applied unchanged to Qwen3.8-27B, matters. The researcher took a published instrument from Anthropic's July workspace paper and asked a simple question: does the tool still work when the model updates? The answer, at least for this architecture and this version step, is a qualified yes. The transferred lens keeps latent entities near the top of the vocabulary on the two-hop reasoning task, and steering directions for "paradox" still suppress the word in Qwen3.8's outputs while keeping the description coherent. That is not a trivial result. It suggests that some interpretability instruments are capturing something more durable than a snapshot of weights.

The practical implications land closer to home than most research updates do. This is not about one model family or one clever trick. It is about whether monitoring pipelines can stop assuming that every release requires a full refit. The researcher is careful about scope: same architecture, same tokenizer, one version step, one lens family. No claims about cross-family transfer or larger gaps. But the numbers are worth sitting with. The median rank for the transferred lens at layer 48 is 17, versus 4 on the home model. That is a degradation, but it is nowhere near the raw logit lens, which sits at ranks of 1e3 to 1e4 through the same band. In other words, the lens is doing real work even when it is not perfectly fitted. And at layer 24, the successor model is actually better than the original, with a median rank of 38 versus 121. The surface readout pays more, and it pays late, but the latent-content readout transfers nearly clean. For anyone building tools that depend on interpretability, that is the difference between a viable product and a research artifact.

This connects directly to the kind of work our publication has been following, like the KV cache as an agent runtime approach, which rethinks what a model's internal state can do during generation, and the EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses research on agents that modify their own execution harnesses. Both of those efforts assume that models are becoming more dynamic, more self-modifying, and more in need of observation tools that do not break the moment a checkpoint changes. This Jacobian lens test is a small but concrete validation of that assumption. It also echoes the spirit of the 44M parameter quantized LLM trained from scratch on 45B tokens, which shows that useful capability can emerge from surprisingly small and efficient setups. The lesson across all of these is that the field is moving toward instruments and models that are built to be interrogated, not just deployed.

The open question is not whether this lens survived. It is whether the pattern holds when the architecture changes, or when the version gap widens from 113 days to something longer. The researcher is honest about the limits of the design, and that honesty is exactly why the result is credible. Our take for anyone building on interpretability tools: do not refit by default. Test your lens on the new checkpoint first. The data here suggests you might save yourself a significant amount of compute and still get a working instrument. But the specific number to watch is the layer-48 transfer cost, which roughly doubles on the surface readout. That is where the next version of this experiment should focus, because if that gap widens predictably with model drift, it gives you a concrete threshold for when a refit becomes necessary. That is a practical detail you can carry into your own work today.

From Machine Learning

Interpretability lenses get fitted to one exact checkpoint, and as far as I can tell nobody had tested what a version update does to one. So this was my question:

when a model line updates, does the fitted instrument survive, or do you refit every release?

Read the original at Machine Learning