AI Agents

Rethinking AI agents on Kubernetes as efficient worker pods

Most teams assume each AI agent deserves its own Pod, but kagent challenges that instinct.

4 min readInfoQ
Rethinking AI agents on Kubernetes as efficient worker pods

There is a quiet assumption running through much of the Kubernetes conversation that if you have an AI agent, you should treat it like a long-running service. Give it a Pod, give it a replica count, and let the scheduler do its thing. The kagent project challenges that reflex by asking a more honest question: what does an agent actually do with compute? The answer, as it turns out, is not much, most of the time. Agents are bursty, they spin up subagents, they pause for human approval, and they often finish in seconds. Reserving an entire Pod for that behavior is like hiring a full-time employee to answer a phone that rings twice a day.

This is where the idea of treating Pods as workers, not agents, starts to make real sense. Instead of forcing a one-to-one mapping between an agent and its execution environment, agent-substrate introduces a control plane that schedules logical "Actors" onto long-lived worker Pods. The Pods stay warm, the agents come and go, and the infrastructure stops pretending that every burst of activity deserves its own dedicated slice of the cluster. It is a more honest model of how these systems actually behave, and it aligns with a broader trend we are seeing across the ecosystem. Consider how Scale Sandboxes Instantly: A New Approach to Concurrent AI Workloads tackles similar waste by rethinking sandbox lifecycle, or how Orchestrate AI Agents: Google Open-Sources AX for Enhanced Efficiency approaches the orchestration problem from a runtime perspective. The through-line is that the industry is finally moving past treating agents like stateless HTTP services.

For our readers, the practical takeaway is not about the specific scheduler implementation. It is about the mental model. When you stop asking "how many Pods do I need for my agents?" and start asking "how many concurrent actors am I actually running at any given moment?" you unlock a completely different cost profile. The kagent argument is not just an efficiency hack; it is a recognition that agent workloads have a fundamentally different shape than traditional microservices. They are conversational, they are hierarchical, and they are often idle. Pretending otherwise means paying for compute that sits unused while your agent waits for a user to click a button or for a subagent to report back.

What we would tell a reader who is evaluating this approach is straightforward: do not adopt this because it is trendy. Adopt it because your agents are already exhibiting these patterns. If you have ever watched a Pod sit at 2% CPU for twenty minutes while an agent waits on an API call, you already understand the problem. The kagent project and agent-substrate are offering a way to stop paying for that idle time, and they are doing it without asking you to abandon Kubernetes or rewrite your entire stack. It is an incremental shift with an outsized impact on utilization.

The open question worth watching is how far this model can be pushed. If logical actors can be scheduled onto shared workers, can the same principle extend to state management, to subagent lifecycles, or even to mixing agent and non-agent workloads on the same Pod? The answer will determine whether this becomes a niche tool or a foundational pattern. Our advice is to experiment with it now, before the pattern is decided for you. The specific detail to watch is whether the control plane can handle the chaotic, branching nature of agent sub-tasks without introducing scheduling bottlenecks of its own. That is where the promise of this approach will either be proven or quietly abandoned.

From InfoQ

Running AI agents on Kubernetes raises a key question: should each agent get its own Pod? The kagent project argues no—agents are bursty, short-lived, can spawn subagents, and may wait for human approval, making one Pod per agent wasteful. Agent-substrate adds a control plane to schedule logical “Actors” onto long-lived worker Pods.

Read the original at InfoQ