Discord's engineering team has done something that most organizations only talk about: they added distributed tracing to Elixir's actor model without breaking performance at scale. That is worth paying attention to, not because Discord is unique, but because their approach reveals a practical truth about observability in modern systems. Tracing every message in a system that handles million-user fanouts is not just expensive, it is impossible. So Discord built a solution that skips the impossible part and focuses on what actually matters.
The key insight from their work is the decision to wrap trace context into a custom Transport library and rely on dynamic sampling. Most teams treat tracing as an afterthought, bolting it onto existing infrastructure and hoping the overhead stays manageable. Discord did the opposite. By filtering context before deserialization and skipping unsampled traces at the CPU level, they recovered more than ten percentage points of overhead. That is not a marginal gain. It is the difference between a system that slows down under load and one that remains transparent to users. For any team running Elixir in production, this is a concrete blueprint for how to make observability work without sacrificing performance.
What makes this story more than a technical case study is the mindset behind it. Discord did not wait for a perfect solution. They accepted that tracing every message is wasteful and designed a system that only pays for what it uses. That is the same kind of pragmatic thinking that separates tools that empower users from tools that get in the way. If you are evaluating how to bring distributed tracing into your own stack, start by asking where the real cost is. Discord's answer is clear: skip the noise, filter early, and measure the overhead before you scale. That advice applies whether you are running Elixir or any other actor-based system.
