Unlocking ASR context with conversation history and prompt-like control.

Automatic Speech Recognition (ASR) models often overlook the potential of prompting, which could significantly enhance their performance in real-world applications.

3 min readMachine Learning

ASR prompting should already be a standard feature in speech models, and the fact that it is not is a genuine obstacle for anyone building voice agents. The developer who posted this observation is doing something straightforward: they are testing whether a speech model can handle a short text prompt, like "Expect a license plate (3 letters, 3 numbers)", and correctly bias its transcription toward that pattern. It works. And yet no major ASR provider offers this capability as a first-class feature.

Consider what that means in practice. Right now, if you run a voice-powered drive-through kiosk and need it to catch license plates reliably, you can use word boosting. You list every plausible three-letter, three-number combination, or you hope the generic boost for "alphanumeric string" is strong enough. It rarely is. You run out of context window, or the boost applies too broadly and starts hallucinating plates where there are none. The developer's alternative, prompting the model with a category, like "Australian cities" or "food names", is more efficient and far more accurate. The model does not need an exhaustive list. It just needs to know what kind of thing to expect.

This is not a futuristic capability. It is a logical extension of how large language models already work, applied to speech. The reason it is missing likely comes down to inertia. ASR systems were built when context meant a list of words, not a natural-language instruction. Existing APIs expose word boosting because that is what the underlying architectures supported. But the developers building voice agents today are not asking for more word lists. They are asking for the model to understand a sentence like "The next utterance will be a person's first and last name" and then transcribe accordingly.

What this means for anyone building voice applications is that the gap between what is possible and what is available is not technical, it is a product design gap. The developer's test shows that a fine-tuned model can handle prompt-like control. The speech model already understands categories. The missing piece is an API that exposes that understanding directly, without requiring users to reverse-engineer a workaround. Until that changes, teams will keep writing brittle boosting lists and hoping the model guesses right. That is not a constraint of the technology. It is a constraint of how we have chosen to talk to it.

From Machine Learning

I've been working on the listening part of my full-duplex speech model and I realized that ASR prompting could be very useful.

Deepgram allows for word boosting but that doesn't work that well in real word applications.

Read the original at Machine Learning