The cost of running a local LLM has always felt like a mystery wrapped in a GPU's thermal throttling. So when someone actually measures the electricity draw of eight models on a single RTX 3090 and reports the price per million tokens in euros, we should pay attention. The headline finding, that the cheapest option wasn't the smallest model and the priciest wasn't the largest, cuts against the lazy assumption that bigger always means more expensive to run. It's a reminder that efficiency isn't a given; it's a design choice, and one that shows up in your power bill long before it shows up in your benchmarks.
This kind of grounded measurement matters because the conversation around AI tools has drifted into the abstract. We're all being told to adopt these systems, but the practical math of what it costs to run one locally, versus relying on a cloud API, remains opaque for most users. The author of the piece did the unglamorous work of plugging in a meter and recording real numbers. That's the same spirit we see in Talking to My AI Clone Taught Me to Question the Tech, where the experience of interacting with an AI clone raises uncomfortable questions about what we're actually building and why. And when we think about the broader shift in Navigating AI/ML Job Requirements: A Shift in Expected Skills, where roles now demand a blend of software engineering and model literacy, understanding the cost structure of running models is no longer a niche concern. It's a core competency.
Our take is that most people are overthinking the hardware arms race. They assume they need the largest possible card or the most recent architecture to get value. This measurement suggests otherwise. The real win isn't in chasing the biggest model; it's in understanding your own workload and matching it to the right tool. The data shows that a mid-sized model can be more cost-effective per token than a smaller one, likely because of how effectively it uses the GPU's memory bandwidth and compute units. That's a practical insight that saves money without sacrificing quality. For anyone considering a local setup, the takeaway is simple: don't trust the spec sheet, trust the meter. And if you're worried about the reliability of what these models are actually outputting, you might want to follow the approach in Verify Your AI's Understanding: A Simple Check for Tax Season, which reminds us that validation is just as important as inference speed.
The open question this raises is whether the industry will start publishing energy costs per model as a standard metric, the way we now see battery life on phones. That would be a genuinely useful shift. For now, the burden is on the user to measure and question. The next time someone tells you a local model is "cheap to run," ask them for the number. If they can't give you one, they haven't checked. And that's the detail worth watching: who will be the first to turn this kind of measurement into a standard feature, not a manual experiment?
