Designing the next generation of AI chips from the inside out

In "Designing AI Chip Software and Hardware," I share my insights from years of experience at Google and Nvidia, where I worked on TPUs and GPUs.

3 min readMachine Learning
Designing the next generation of AI chips from the inside out
[R] Designing AI Chip Software and Hardware

The most practical thing about this document is that it exists at all. A former Google TPU and Nvidia GPU engineer publishes a full blueprint for a competitor's AI chip, complete with software architecture, hardware trade-offs, and career anecdotes, all because he decided not to start the company. That is not a leak. It is a gift to anyone building in this space, and it deserves to be read as one.

What makes this more than a technical curiosity is the willingness to show the work. He does not claim his design is better than TPUs or GPUs. He simply walks through the decisions he would have made, the constraints he would have prioritized, and the reasons behind each choice. For engineers and product leaders evaluating their own hardware strategies, this is rare signal. Most chip design discussions stay locked inside NDA-bound rooms or emerge years later as polished case studies that omit the hard parts. Here we get the hard parts, including the dead ends, the personal motivations, and the candid admission that he ultimately walked away. That honesty creates trust.

For readers who are not hardware designers, the value is different but real. The document forces a question that too many teams avoid: what would you build if you started from the user's workflow instead of the component spec? The approach treats software and hardware as a single system, not two teams throwing designs over a wall. That is the kind of thinking that makes AI tools feel intuitive rather than powerful-but-frustrating. It is also the kind of thinking that legacy spreadsheet vendors never seem to apply to their own architectures, which is why users still fight with formula errors and slow recalculations.

The practical takeaway for our readers is straightforward. If you are evaluating AI hardware for your own work, ask whether the vendor thinks this way. Does the chip design start with the actual task a user wants to accomplish, or does it start with a spec sheet full of teraflops and memory bandwidth? The second approach produces benchmarks that look great in a press release and feel terrible in daily use. The first approach, the one this document exemplifies, builds tools that fade into the background and let you focus on the problem you actually care about. That is the only kind of AI chip worth waiting for.

From Machine Learning

This is a detailed document on how to design an AI chip, both software and hardware.

I used to work at Google on TPUs and at Nvidia on GPUs, so I have some idea about this, though the design I suggest is not the same as TPUs or GPUs.

Read the original at Machine Learning