Atlassian's recent migration from gostatsd to OpenTelemetry is the kind of infrastructure story that rarely gets told, but it matters more than most product launches. The team replaced the internal engine of a metrics platform ingesting data from roughly 100,000 hosts across 14 regions, all while maintaining a 99.95% service-level objective. That is not a weekend project. It is a careful, deliberate surgical procedure on a system that production alerts depend on every second. For anyone who has ever hesitated to touch a legacy pipeline for fear of breaking downstream monitoring, this account is worth reading closely.
What makes this story instructive is not the technology choice itself, OpenTelemetry is becoming the standard for observability data collection, but the execution philosophy. Atlassian did not rip and replace. They changed the internal workings without disrupting the metrics that teams had already built alerts around. That constraint forced them to think about compatibility, data continuity, and testing in ways that a greenfield deployment never would. It echoes a pattern we have seen elsewhere: Self-Service GPU Metrics Bring Team Clarity Without Cross-Team Exposure describes how Adobe gave teams independent access to Prometheus metrics without opening up broader infrastructure. Both stories are about enabling better observability without creating new risks or dependencies. The difference is that Atlassian had to preserve existing alerting logic exactly, which is a harder constraint than building new access paths.
The practical takeaway here is that migrating to OpenTelemetry does not have to mean rewriting your alerting rules or retraining your teams on new dashboards. If you are considering a similar move, the question to ask is not "Can OpenTelemetry handle our scale?" but "Can we replace the plumbing without the users noticing?" Atlassian has shown that the answer is yes, provided you invest in rigorous testing and incremental rollouts. That is a concrete, quotable insight: migrate the pipeline, not the policies. Your alert thresholds, your SLOs, your team's mental model of what a metric means, all of that can stay intact while the underlying collection layer modernizes.
One detail worth watching is how Atlassian handled the inevitable edge cases that emerge when two different metric-collection systems run side-by-side. The article notes that the existing service had a 99.95% SLO, meaning that even a brief data gap or a subtle change in metric semantics could trigger false alerts or missed detections. The team's approach to validating that the new pipeline produced identical outputs to the old one is the kind of engineering discipline that separates a successful migration from a post-mortem. For any organization running hundreds of thousands of hosts, the next frontier is not just adopting OpenTelemetry, it is proving that the new system is indistinguishable from the old one at the data level, before you turn off the legacy service. That is the detail that will determine whether your migration is a quiet success or a noisy incident.