Most conversations about machine learning performance start and end with the same number: FLOPs. Reduce the floating-point operations, the thinking goes, and the model runs faster. The engineer behind *How to Make Your Model Fast: A Systems View of Efficient Machine Learning, from Silicon to Agents* spent months writing the guide he wishes he had when starting out, and his central insight is worth repeating: reducing FLOPs does not guarantee speed. What matters is understanding what the system is actually bounded by. This kind of systems-level thinking is rare in a field that often rewards the flashiest optimization trick. It is also exactly what teams need as they move from training models to running them in production, especially as agentic workflows add new layers of complexity. We have seen similar themes emerge in recent coverage, such as how Self-Service GPU Metrics Bring Team Clarity Without Cross-Team Exposure helps engineers identify where their throughput is stalling, and how Explore Open-Source Tools That Give AI Agents Lasting Memory addresses a different kind of bottleneck, context persistence across sessions. Both point to the same truth: optimization starts with diagnosis, not guesswork.
The guide's structure reveals a clear progression from the hardware up, starting with roofline analysis and moving through kernels, compilers, quantization, pruning, on-device LLMs, and finally agents. That ordering matters. Too many practitioners jump straight to quantization or kernel fusion without first asking whether the model is compute-bound, bandwidth-bound, memory-bound, or system-bound. The author frames the problem as a set of diagnostic questions: How fast can this model possibly run on this hardware? Which optimization will actually move the limit? Is quantization even worth doing here? These are not rhetorical questions. They require a methodical approach, and the open-source format invites contribution and critique from the community. For anyone building on-device models or serving real-time inference, this is a practical resource, not a theoretical treatise. It also connects naturally to the work being done on Secure AI agents with new guardrails for safer autonomy, because a system that is poorly optimized is also harder to secure, latency and throughput constraints often force tradeoffs that compromise safety.
Our take is straightforward: this guide earns its place on your bookmarks because it resists the temptation to sell a single silver bullet. The author is explicit that the same systems thinking applied to a single model can and should be extended to serving infrastructure and agent systems. That is the detail to watch. As agents become more autonomous, the performance bottlenecks shift from raw compute to orchestration overhead, memory management, and inter-agent communication. The principles in this guide, understand your bounds, measure before you optimize, reason from hardware up, apply just as forcefully to a multi-agent pipeline as they do to a single transformer layer. The concrete takeaway for our readers: before you run another optimization pass, ask what the system is bounded by. If you cannot answer that question, the optimization is a gamble. This guide gives you the tools to stop guessing.