Latent Reasoning

Beyond the Token Stream: Architectures for Deeper Reasoning

The field is quietly shifting its focus from longer chains of thought to something more fundamental: reasoning that never touches language at all.

4 min readMachine Learning

The conversation around latent reasoning has reached a tipping point, and the post you are referencing captures why. The observation that large language models routinely generate flawless chains of thought that end in wrong answers, or correct answers with flawed logic, is not just an academic curiosity. It is the clearest signal we have that the token stream is a distraction. The trace is not the computation. If the field is serious about moving toward more general intelligence, it has to stop optimizing for the appearance of reasoning and start building systems where the reasoning is the hidden state itself. This is the same practical impulse driving work on Unlock LLM Training: A Practical Guide to Distributed Algorithms, where the focus is on the mechanics of scale rather than the theater of output. The architecture matters more than the monologue.

Mapping the five families of latent reasoning is a useful discipline because it forces a clear separation of concerns. On one end, you have Coconut and Soft Thinking, which feed hidden states back into the model and argue for parallel search frontiers compressed into a single continuous representation. On the other, you have Abstract-CoT, which trades verbal rationales for a learned non-linguistic vocabulary, still serial but no longer human-readable. In between sit the recurrent and recursive solvers, from looped Transformers to task-trained systems like HRM and TRM, which refine latent states through repeated application. The distinction that matters most is not just where the computation happens but how a system acquires a new task. Whether through context, memory, or gradient updates changes the economics of adaptation entirely. The BDH-CQ approach, built on the Dragon hatchling architecture, is interesting because it writes demonstrations directly into a recurrent memory at inference time, solving unseen tasks without a backward pass. That is a fundamentally different trade-off from the transductive ARC pipelines that require optimization per puzzle.

For practitioners, the immediate question is not which family wins, but what we are willing to give up. The industry has built a significant portion of its interpretability and evaluation toolkit on the assumption that readable traces are a window into model reasoning. That assumption is already shaky, as the Kambhampati reference makes clear. If latent reasoning becomes the default, we lose the ability to audit intermediate steps, and we will need new tools to build trust. This is not a hypothetical concern. It is the same tension that surfaces when you evaluate models for practical decision-making rather than benchmark performance, a theme explored in Jev vs LLMs: Evaluating AI for Practical Decision-Making. Accuracy on a test set tells you little about the soundness of the path taken to get there. The push toward edge deployment and open-weight models, as seen in Unlock AI on Your Glasses: PrismML's Innovation Powers Smarter Devices, only intensifies the need for inference that works within tight compute budgets, where verbose CoT is a luxury we cannot afford.

The open question worth watching is whether CoT legibility was a temporary artifact of scaling or a genuine safety property. If the latter, we are heading toward a future where we trade transparency for efficiency and call it progress. The early pretraining results from BDH-CQ, showing transformer-like scaling laws up to 600B parameters while preserving latent reasoning behavior, suggest the efficiency gains are real. But efficiency for what? If we cannot inspect why a model made a decision, we will spend the next decade building interpretability tools for systems that were not designed with interpretability in mind. The concrete point to watch is whether the community starts treating latent traceability as a first-class evaluation metric, alongside accuracy and cost. That would be a meaningful shift. Until then, the field is betting that the hidden state knows best, and for the first time, the evidence might actually support that bet.

From Machine Learning

After following various arXiv papers and researcher discussions on X/bluesky about latent reasoning and continual learning, one idea which resonates strongly is that path forward (towards AGI) may depend less on generating ever-longer chains of thought and more on finding architectures that can reason beyond the token stream.

LLMs routinely reach correct answers through flawed or fabricated CoT steps, and produce perfectly logical steps that end in wrong answers (Kambhampati, 2025). The trace doesn't track the computation which clarifies that verbalized CoT is an imitation of reasoning and not the mechanism itself.

Read the original at Machine Learning