Transformer

From Paper to Practice: Training a Transformer for English-Tamil Translation

Building a Transformer from scratch in pure PyTorch is no small feat.

4 min readMachine Learning

There is something quietly radical about building a Transformer from scratch in pure PyTorch, and not because the architecture is new. The original "Attention Is All You Need" paper is nearly a decade old, and the field has moved on to massive pretrained models that most of us interact with only through APIs. What Imran Coders has done with an English-to-Tamil translation model is strip away the abstraction layer that usually sits between us and the math. He trained on dual NVIDIA T4 GPUs using a Hugging Face dataset, and he published the full breakdown, equations, tensor shapes, and all. That is not just a tutorial. It is a map of how the machinery actually works.

We see a lot of content about the bleeding edge of large language models, and much of it assumes you already understand the underlying mechanics. But there is a growing gap between people who can prompt a model and people who can build one. This project is a direct challenge to that gap. It is one thing to read about attention mechanisms or to watch a video on positional encodings. It is another to sit with the code and watch every dimension transform through each layer. For anyone who has felt the frustration of hitting a wall with a high-level framework, this kind of work is the antidote. It is also worth noting the choice of language pair. English-to-Tamil is not a toy dataset. It is a real translation task with its own structural complexities, and it forces the builder to engage with more than just token arithmetic.

This approach connects to a broader conversation we have been following, particularly around how Unlock LLM Training: A Practical Guide to Distributed Algorithms and Exploring Paragraph Structure: How LLMs Navigate Token Space both zero in on the mechanics that most users never see. The former tackles the systems engineering side, the latter looks at how token positions create meaning. Both are useful, but neither replaces the experience of training your own model from scratch. And while Scale AWS Server Deployments Effortlessly with Stateless Model Context Protocol is about serving at scale, the same principle applies: you cannot debug what you do not understand.

Our honest take is that this is the kind of work we should point more people toward. Not because it is flashy, but because it is foundational. The author has done the hard part, which is not the code itself but the willingness to explain every step. The math is dense, and the tutorial is not a light read. But that is the point. The barrier to entry in this field is not compute or even data. It is the courage to open the hood and see how the engine runs. We would tell anyone who asks us about this to start with the repository, run the training loop, and then change one thing, just one, and see what happens. The real lesson is not in the Transformer architecture. It is in the confidence you build when you realize you can trace every number back to its source.

The open question we are left with is simple: what comes next? A model trained on two GPUs is a proof of concept, not a production system. But the point was never to ship a translation service. The point was to demonstrate that the underlying technology is understandable, and by extension, improvable. If more people did this, the gap between prompters and builders would narrow. That would be a future worth exploring.

From Machine Learning

I built and trained the complete Transformer architecture from scratch using pure PyTorch (`torch.nn` primitives) based on the original "Attention Is All You Need" paper.

I trained the model on an English-to-Tamil parallel translation dataset (`gopi30/english-tamil` on Hugging Face) using dual NVIDIA T4 GPUs on Kaggle.

Read the original at Machine Learning