Transformer

From Python to weights: a compiler proves what transformers can express

Most people assume a transformer's capabilities come from training.

4 min readMachine Learning

Most people assume a transformer's intelligence lives in its training. Feed it enough data, adjust the weights, and eventually it learns to reason. That story is true, but it is incomplete. There is a quieter question underneath: what can a transformer actually express, regardless of how it got those weights? This week, a developer going by notforrob posted a compelling answer. They built a compiler that takes a computation graph written in ordinary Python and outputs the weights of a standard Phi-3-architecture transformer. No training, no gradient descent, no trust_remote_code. Vanilla Hugging Face loads it and runs it. The write-up and twelve runnable examples are public, and the implications reach far beyond a clever hack.

This is not the first time someone has hand-crafted transformer weights. RASP gave us a language for describing attention and MLP primitives, and Tracr compiled those programs into actual weights. What stands out here is the target. Tracr and its descendants often produce weights that work in principle but require custom scaffolding or modified architectures. This project deliberately targets a stock, off-the-shelf model architecture. That choice matters because it collapses the distance between "a transformer can execute this algorithm" and "this exact model you already know how to run can execute this algorithm." It reframes the model not as a black box that learned a behavior, but as a programmable machine whose weights are a compiled artifact. If you have been following how LLMs navigate token space, you already understand that where a token sits in the sequence is a coordinate. This compiler treats that coordinate system as something you can deliberately construct rather than merely observe.

For the rest of us, the practical takeaway is not about writing a compiler. It is about what this reveals regarding the boundary between architecture and training. We tend to treat model weights as the product of data alone. This work shows that a nontrivial class of algorithms can be embedded directly into a standard architecture by construction. That does not make training obsolete, but it does suggest that some capabilities do not need to be learned from scratch. They can be designed in. If you are building on top of LLMs, this changes how you debug and reason about failure modes. When a model fails at a task, the usual suspects are data quality or model size. This project suggests another possibility: the architecture might be expressing an algorithm you did not realize it was running, and now you have a way to test that hypothesis directly. The related discussion on distributed training algorithms often focuses on scale, but this is a different axis of control, one that is about deliberate construction rather than sheer volume.

What makes this worth watching is not the novelty of hand-built weights. It is the discipline of targeting a stock architecture. That choice forces the compiler to solve the hard problem of mapping abstract operations onto the exact tensor shapes, attention patterns, and feedforward blocks that a real model expects. It is a stress test for our understanding of what each component in a transformer actually computes. The open question I would put to anyone reading this: if you can compile a computation graph into a vanilla checkpoint, what stops the reverse process? What if, instead of asking what a trained model learned, you could decompile its weights back into a human-readable program? That is the direction this points toward. And for now, the concrete thing to watch is whether the community builds on this to create a library of verified, hand-crafted behaviors, one weight matrix at a time.

From Machine Learning

I've been chasing the question of what algorithms a transformer can actually express -- separate from what it can learn. So I built a compiler: define a computation graph in ordinary Python, and it produces the weights of a transformer that executes the graph. The result is a standard Phi-3-architecture checkpoint that vanilla huggingface loads with no custom code and no trust_remote_code. Zero training in the pipeline.

Write-up (origin + how the constructions work): https://ood.dev/posts/torchwright-intro/

Read the original at Machine Learning