Transformers

Teaching a transformer exact arithmetic by hand, not training

A single transformer model, its weights hand-set with no training, just averted the arithmetic meltdown that defines its peers.

4 min readMachine Learning

Somewhere between the benchmark arms race and the endless parade of "frontier" model announcements, a quiet act of defiance is happening. A developer, not a lab, took a stock Phi-3 transformer and refused to train it. Instead, they reached into the weights and manually compiled the grade-school multiplication algorithm directly into the parameters. No gradient descent. No fine-tuning. Just a computation graph translated into an ordinary Hugging Face checkpoint via a custom compiler called Torchwright. The result: 100% accuracy across all three million supported three-digit expressions, with checkpoints scaling up to 12-digit by 12-digit multiplication. That is not a marginal improvement. That is a statement about where the real leverage in AI lies.

For anyone who has watched the industry chase scale like a fixed destination, this is the reality check worth sitting with. The developer was explicit that nobody needs a transformer for arithmetic, and they are right. But the point is not the multiplication. The point is that we have spent years treating training as the only path to capability, when this work demonstrates a different lever entirely: direct weight manipulation. It is the difference between teaching by rote and writing the answer key in permanent ink. While frontier models famously fall off a cliff at seven digits, scoring 0/500 across five of six tested systems, this compiled transformer stays perfect. Yes, the advantage is that the algorithm was literally embedded. But that is exactly the lesson. We have been so focused on making models larger that we have overlooked how much control we can already exert over their internals.

This connects to a broader shift we have been tracking across the ecosystem. When Compile TypeScript to Native Code and Transform Your App Performance landed, the takeaway was similar: performance is not just about better hardware, it is about better compilation. And in the context of Scale AWS Server Deployments Effortlessly with Stateless Model Context Protocol, we see the same instinct applied to infrastructure: strip away unnecessary complexity and state to gain efficiency. Torchwright fits squarely into that lineage. It treats the transformer not as a black box to be coaxed, but as a programmable machine with a known architecture. That is a fundamentally different posture from the "throw more data at it" school of thought.

What we would tell a reader who asks whether this matters: yes, but not for the reason you think. The concrete takeaway is that the boundary between traditional software and neural networks is dissolving. If you can compile an algorithm into weights, then the model is no longer just a statistical approximation; it is a deterministic engine with a known behavior. That has implications for verification, for safety, and for debugging. We can finally ask questions like "what exact function is this model computing?" and get a precise answer. The open question is whether this approach scales beyond arithmetic. If it can be generalized to other discrete operations, or to hybrid models that mix learned and compiled components, we are looking at a new class of interpretable AI. That is the detail to watch. Not the next billion-parameter release, but the quiet compiler that lets us write the rules back into the machine.

From Machine Learning

Obviously nobody needs a transformer that's good at multiplication. I wanted to know whether a stock transformer could do exact arithmetic if I chose its weights directly.

I implemented the grade-school algorithm as a computation graph and compiled it into an ordinary Phi-3 Hugging Face checkpoint using Torchwright, a compiler I wrote. No training. The three-digit calculator gets all 3,000,000 supported expressions right. I've published checkpoints to Hugging Face that support up to 12 digit x 12 digit multiplication.

Read the original at Machine Learning