DoorDash's recent work building a GenAI platform for over 5,000 internal users is the kind of engineering story that rewards close reading. Swaroop Chitlur and Sidd Kodwani walk through a shift that many organizations will soon face: moving from vendor-first setups to open-weights models, and building the gateways and agent infrastructure to make both work at scale. What stands out is not the size of the deployment, it's the architectural discipline behind it. This is not a story about hype. It is a practical case study in balancing accuracy, latency, and cost when the tools themselves are evolving weekly.
For our readers, the most actionable insight here is how DoorDash treated model selection as an operational decision, not a philosophical one. They started with vendor APIs, then migrated to open-weights models as their use cases matured. That progression mirrors what we see in smarter AI deployments across the industry: the best stack is the one you can change. If you are building an internal platform today, the goal should not be to pick the "best" model once. It should be to build the routing and evaluation layers that let you swap models without rebuilding everything. This connects directly to what we explored in a recent piece on Grant Your LLM Safe Autonomy in 9 Practical Steps, where the focus was on giving models permission to act without losing control. DoorDash's agent gateway, a dedicated layer for LLM orchestration, is exactly that kind of infrastructure in practice.
The implications for teams building smaller-scale systems are worth spelling out. You do not need 5,000 users to learn from DoorDash's approach. Their emphasis on separating LLM gateways from agent gateways is a design pattern any team can adopt. The LLM gateway handles model routing, caching, and cost tracking. The agent gateway manages tool access, memory, and task orchestration. Keeping them distinct prevents the common mistake of conflating model capability with execution safety. This also echoes findings from a recent study we covered on how AI models bend facts for verified sources, a reminder that even when the model performs well, the system around it must enforce factuality and safety boundaries. DoorDash's architecture does exactly that, not by restricting users, but by making the guardrails transparent and tunable.
One specific detail from the DoorDash build is worth watching closely: how they handle accuracy trade-offs. With 5,000 internal users, the platform cannot afford a one-size-fits-all threshold. Different teams, logistics, finance, support, have different tolerance for error. DoorDash's solution was to expose configurability in the gateway layer, letting teams set their own accuracy-cost preferences per use case. That is a design choice that turns a platform from a bottleneck into an enabler. For anyone building internal GenAI tools, the question is not whether your model is accurate enough. It is whether your platform lets each team decide what "enough" means, and whether you can measure the gap. Until you build that feedback loop, you are guessing. DoorDash stopped guessing.
