There's a number in this story that should stop you cold: at Kilo Code, engineers are reading or writing code themselves about one percent of the time. One percent. The rest is agents. That is not a productivity hack or a nice efficiency gain; that is a fundamental reordering of what a software engineer actually does. It also means the old arguments about AI assistance are over. The question is no longer whether agents should touch your codebase, but which systems you trust them to touch at all. As Scale AI Workflows: Modernizing APIs with Architecture as Code shows, even massive enterprises are rethinking how they structure the very foundations of their code. The difference is that Morgan Stanley is doing it deliberately, with an eye toward architecture; Kilo Code and its peers are doing it because they have no other choice.
The real tension here is between greenfield and brownfield. Symbotic's Jared Go puts it plainly: agents are great at building new codebases, but updating and maintaining existing ones is where the actual challenge lies. That is the honest version of the story, and it is the one most vendors would rather not talk about. Everyone wants to demo the agent that spins up a fresh service from a prompt. Nobody wants to talk about the agent that has to untangle a legacy C# monolith at three in the morning. Replit's approach offers a more practical middle ground: they let agents do the work, but an agent reviews every pull request and assigns a risk score. Low-risk changes get merged automatically; anything with a pulse goes to a human. That is not surrender to the machines, and it is not Luddite resistance. It is the only sensible way to run a software organization when the cost of a bad merge is measured in downtime, not in tokens. For teams feeling the pressure to go full-agentic, that distinction is the difference between a tool and a liability.
The cost conversation is where this story gets genuinely uncomfortable, and it is also where we would push back on the triumphalism. Emilie Schario is tracking cost per pull request, which is a smart metric, but she also admits she has an engineer with a heavy foot who is constantly at the top of the usage board. Replit found a support user who blew through an insane amount of money running an automation on a frontier model. The details are different, but the pattern is the same: the technology works, so nobody questions the bill until it arrives. That is not a sustainable way to run an AI program, and it is not something you can fix with better model routing alone. You need visibility that is not anti-productive, as Replit's Amol Jain puts it. You need sensible defaults, because most tasks do not need the frontier. And you need to accept that the answer to runaway costs is not to stop using the tools; it is to get serious about governance. As Automating Feature Flag Cleanup: DoorDash Leverages AI for Efficiency demonstrates, the teams that succeed with AI are the ones that treat it as a managed system, not a magic trick.
Here is the takeaway we would give any engineering leader reading this: the era of asking whether AI can write your code is over. It can. The question you need to answer is whether you have the operational backbone to survive the answer. That means building risk-scoring into your review process, adopting multi-model routing so you are not paying frontier prices for CRUD operations, and most importantly, accepting that human oversight is not a bottleneck. It is a feature. Replit's story about the gnarly bug that took six hours and a fleet of agents to fix is impressive, but do not miss the detail that the AI manager agent had to tell another agent to go to sleep first. The tools are powerful, but they still need a human hand on the wheel. The teams that figure out how to provide that without slowing down are the ones that will win. Everyone else is just burning tokens.
