Whisper

Discover how AI models run entirely offline on your iPhone

Running Whisper, Qwen3-ASR, Nemotron, and MOSS entirely offline on an iPhone is no small feat.

4 min readMachine Learning

The quiet ambition of LiveTranscriber is easy to miss if you skim past the model names. Over the past month, a developer going by marshmallow_ki built an open-source iOS app that runs Whisper, Qwen3-ASR, NVIDIA Nemotron Streaming, and MOSS entirely on-device. No cloud, no latency hiding in a server farm, no privacy trade-off. The stated goal was to see if recent open-source models could become a practical mobile product, not just a benchmark slide. That question matters more than the app itself, because the answer reveals where mobile AI is heading for the rest of us.

The engineering story here is not about squeezing a neural network onto a phone for the sake of a demo. It is about memory management, streaming latency, model loading, context handling, battery drain, and switching between inference backends. That is the unglamorous work that separates a parlor trick from a tool you actually open daily. We have seen this pattern before in adjacent fields. For example, Bridging Embedded Systems Expertise to the World of Machine Learning shows how developers with low-level systems backgrounds are often the ones who make AI feel tangible, because they understand the constraints that cloud-first engineers can ignore. LiveTranscriber is the same instinct applied to your pocket: it is not enough that a model exists; it has to run within the physical limits of a phone while still feeling responsive.

What makes this worth your attention is not the feature list, though offline transcription, speaker-aware recognition, and on-device summaries are genuinely useful. The real signal is the shift in what is possible without a network connection. Most of us have accepted that speech-to-text and summarization require sending audio or text somewhere else. This app challenges that assumption directly. It also connects to a broader tension we have been tracking: the gap between what AI can do and what it costs to run. When you read about The KV Cache Tax: Why Inference Servers Run Out of Memory Before Compute, you see that even data centers struggle with the memory demands of modern models. LiveTranscriber does not sidestep that problem; it confronts it on a device with a fraction of the resources. That is not a minor achievement. It is a proof point that the frontier is moving toward efficiency, not just raw scale.

Our take is straightforward: this is the kind of project that should make you rethink what you expect from your phone. If a solo developer can turn four serious models into a usable iOS app in a month, the bottleneck is no longer capability. It is design and distribution. The open-source nature of the project means others can learn from the backends, the batching logic, and the trade-offs made for battery life. We would tell a reader who asks whether this matters to them: watch how quickly similar capabilities become standard in the apps you already use. The specific models will change, but the pattern of local, private, low-latency AI is here to stay. The concrete detail to watch is whether Apple's own on-device frameworks start absorbing these lessons, or whether third-party apps like this one continue to lead the way by a comfortable margin. That answer will determine if your next phone feels like a tool or just a receiver.

From Machine Learning

Over the past month, I've been building LiveTranscriber, an open-source iOS app for running modern speech and language models entirely on-device.

The goal was to see whether recent open-source models could be turned into a practical mobile product—not just technical demos.

Read the original at Machine Learning