Continuous Delivery

Progressive Deployments for Stateful Infrastructure That Actually Works

Core infrastructure fails differently than applications, and pretending otherwise is how outages happen.

3 min readInfoQ
Progressive Deployments for Stateful Infrastructure That Actually Works

Most engineering teams treat CI/CD as a solved problem until they hit the wall of stateful infrastructure. Ian Nowland's recent talk on continuous delivery for foundational platforms names that wall explicitly: conventional pipelines assume stateless services, but core platforms carry data, state, and consequences that a rolling deploy can't simply shrug off. That's the gap between deploying a service and evolving a system that others depend on. Nowland, who led reliability work at AWS and Datadog, isn't arguing for abandoning CI/CD. He's arguing for a more honest version of it, one that accepts that production is the only environment that truly counts.

The practical value here is that Nowland isn't selling a tool or a dashboard. He's describing a discipline: progressive deployment, synthetic testing in production, and deliberately shrinking the blast radius before you touch a single node. These aren't new ideas, but they're rarely treated as a coherent practice for foundational platforms. Most teams treat canary releases as an afterthought or a checkbox. Nowland's framing suggests they should be the default design constraint from the first commit. That resonates with the broader shift we're seeing in how platforms are built and operated. For example, Scale AWS Server Deployments Effortlessly with Stateless Model Context Protocol shows how even stateless patterns are being rethought for operational simplicity, while Monitor Cypress Tests with Grafana: Persistent Observability for Your Data reinforces that observability isn't a side effect, it's a requirement for knowing what your deployments actually did.

What we'd tell a reader who asks "where do I start?" is this: stop optimizing your pipeline's speed and start optimizing for your ability to recover. Nowland's emphasis on synthetic testing in production is the clearest signal of that mindset. You're not testing to prove something works; you're testing to catch what breaks before it reaches a user. That flips the usual logic of pre-production QA. And it aligns with the conversation in Unlock AI’s Enterprise Potential: Navigating Adoption and Ethical Considerations about how enterprises adopt new capabilities responsibly, not because they're forced to, but because they've built the muscle for controlled risk.

Our honest take is that most teams don't have a deployment problem. They have a trust problem with their own infrastructure. They don't believe their tests, they don't trust their rollback, and they're terrified of the blast radius because they've never measured it. Nowland's talk is useful precisely because it treats those fears as engineering problems rather than personality flaws. The specific takeaway worth quoting: "If you can't shrink the blast radius, you're not ready to deploy faster." That's the line to bring back to your team. The open question is whether your platform team will treat that as a constraint or an excuse.

From InfoQ

Ian Nowland discusses why conventional CI/CD practices break down for stateful, core infrastructure. Drawing from his leadership at AWS and Datadog, he shares actionable techniques for safe progressive deployments, synthetic testing in production, and mitigating blast radius in complex software platforms.

Read the original at InfoQ