Distilling Smarter AI Assistants from Real User Feedback

Every time a user clicks "accept" on an AI-generated fix, they're casting a vote.

3 min readInfoQ
Distilling Smarter AI Assistants from Real User Feedback

Most teams treat AI code assistants as black boxes. You prompt, you accept or reject, and the system learns nothing from that choice. Ben O'Mahony's presentation flips that dynamic on its head by treating your cursor as a telemetry device. Instead of waiting for explicit feedback, he instruments the agent itself with OpenTelemetry to capture implicit labels: did you accept that fix, dismiss it, or ask for a regeneration? That simple triad of actions becomes the training signal for a smaller, local model that distills what a frontier model does well without the per-token cost or privacy risk.

This is the first time we've seen a practical bridge between production observability and model distillation aimed at the developer tooling layer. The OpenTelemetry ecosystem has long been about tracing requests across services; here, it's tracing decisions across a human-AI loop. That matters because implicit labels are noisy but abundant. You don't need a human to write a paragraph explaining why a suggestion was wrong. You just need them to hit "dismiss" a thousand times. Over that volume, patterns emerge, and those patterns are exactly what a small language model can learn to replicate. It's not about copying the frontier model's reasoning; it's about copying the user's judgment of that reasoning.

For our readers, the practical takeaway is direct: you already have the data to build a better internal assistant, and you don't need a research lab to start. If your team runs a custom Language Server Protocol today, you can add OpenTelemetry instrumentation to capture accept and reject events on every hover, completion, and quick fix. That telemetry becomes a continuous data flywheel, not a one-off dataset. The frontier model stays in the loop during development, but over time, the local SLM learns which suggestions your team actually trusts. That is a cost play, a latency play, and a privacy play all at once. We would tell any team building an internal copilot: stop guessing what works, start measuring what your engineers choose to accept.

The open question is whether implicit labels are enough to capture the nuance of "good" code, as opposed to merely "accepted" code. Acceptance can be lazy. A developer might accept a mediocre fix because it's faster than typing the alternative. O'Mahony's approach doesn't solve that bias, but it does surface it. If you see a high acceptance rate on a specific pattern, you can dig into whether that pattern is genuinely good or just path-of-least-resistance. That audit trail is a feature, not a bug. The concrete detail to watch is how your team's dismissal patterns change over time, because that is the earliest signal of whether your distilled model is drifting toward compliance over quality. If you're not tracking that, you're flying blind with a faster autopilot.

From InfoQ

Ben O'Mahony discusses building custom AI-powered Language Server Protocols (LSPs) that go beyond standard rule-based checkers. He explains how to instrument AI agents natively with OpenTelemetry to track concrete user actions (accepting, dismissing, or regenerating code fixes) as implicit labels, creating a continuous data flywheel to distill frontier capabilities into cheaper, local SLMs.

Read the original at InfoQ