For years, the story of developer productivity has been told in the editor. We measure lines of code, pull request velocity, and commit frequency, as if the act of typing is where value lives. GM's autonomous driving division just published a counter-argument that should reshape how you think about your own toolchain, and it has nothing to do with writing faster. Their engineers spend only 15% of their time writing code. The other 85% is analysis, triage, testing, and the invisible glue that turns a hunch into a working system. If you are still optimizing for the 15%, you are leaving the real bottlenecks untouched, a lesson that applies well beyond the automotive industry. This is the same logic that drives our Clean Data Starts With Catching AI Slop Before It Skews Your Model, where the bottleneck wasn't model architecture but the garbage flowing into it.
The headline number is triple the merged pull requests, and that is worth pausing on. But what impresses us is not the volume; it is the diagnosis. GM did not hand every developer a chatbot and hope for the best. They mapped the full engineering loop, from problem discovery in vehicle telemetry to a verified fix in production, and then asked where time actually disappeared. They found that agents could analyze petabytes of road data, perform initial triage, and even run machine learning experiments in parallel. Crucially, they gave those agents access to internal tools through Model Context Protocol servers, not just a chat window. The result is a workflow where an agent can identify a potential issue, locate the affected component, search historical data for similar incidents, and present a human-readable case to an engineer. That is not automation for its own sake; it is a deliberate restructuring of how work flows. For anyone building with AI, this is the difference between a novelty and a system. It is the same distinction we explored in Exploring Real-World Computer Vision: Deployments, Edge Models, and Current Challenges, where the real challenges were not in the model card but in the messy, real-world deployment.
Our take is simple: GM is treating the agent as a colleague with a scoped job, not a magic wand. They assigned four deployed engineers to work directly with teams, helping them identify useful workflows and spread successful practices. They based an agent's permissions on the engineer using it, and they kept the human accountable for the output. That is a governance model we would bet on, because it addresses the real reason most AI initiatives stall: trust. You cannot trust a system you do not understand, and you cannot understand it if it operates as a black box. GM's insistence on human-readable outputs and version-controlled skills is the antidote to the black box. The fact that they saw fewer escaped defects suggests this is not just faster; it is safer. The takeaway you can quote: "If you give somebody just a chatbot which can do coding, there is still a lot of inefficiency built into that process." GM's real innovation was removing the inefficiency around the code, not in the code itself.
The open question for you is not whether to adopt agents, but where to draw the first loop. GM started with the longest bottleneck in each cycle, not the easiest task. They automated the analysis and triage first, leaving the final verification to a human. That is a concrete, repeatable pattern. The next time you look at your own pipeline, ask yourself what consumes the 85% of time that is not typing. If you can map that, you can automate it. If you cannot, you are not ready for agents. We will be watching to see if GM's model survives contact with the chaos of production, but for now, they have given us a rare gift: a clear, honest blueprint for what agentic AI actually looks like when it is built for the work, not for the demo.
