What on-device means for speech AI
A speech recognition system has two broad parts: preparing the audio, and running a trained model that turns sound into text. On-device recognition means the model itself runs locally, on the device's CPU, GPU or a dedicated neural processing unit. Preparing audio locally while sending it to a server for recognition is a hybrid design, not on-device recognition, even though the preparation happens on your machine.
That distinction matters because marketing often blurs it. A tool can process your video locally, compress the audio locally, and still send that audio to a server for the actual transcription. The broader architecture question is covered in client-side vs server-side video processing; this article focuses on the speech model itself.
What already runs locally
- Phone operating systems offer on-device dictation for some languages, often as an option or as the default when the language model has been downloaded.
- Live captions features on some phones and computers run recognition locally for media playing on the device.
- Open models such as Whisper come in several sizes, and community ports run the smaller versions on ordinary laptops and some phones.
- Browsers increasingly expose hardware acceleration to web pages, which makes in-browser models possible, though performance varies widely by device.
The pattern is clear: short-form, real-time speech tasks in major languages are increasingly local. Long recordings, many languages and the highest accuracy are where cloud processing still tends to win.
Privacy and latency: the two main reasons to go local
Privacy is the headline benefit. If recognition happens on the device, the audio doesn't travel to a provider, isn't stored on someone else's server and can't be exposed in a breach there. Local processing doesn't protect against everything, though; a compromised device, a synced backup or a shared transcript still leak. The limits are discussed in on-device video privacy, and the questions to ask a cloud provider are in cloud speech processing privacy.
Latency is the second benefit. Without a round trip to a server, dictation and live captions can respond faster, and they keep working offline, on a plane or in a basement with no signal. For a finished recording that you are happy to wait a few minutes for, latency matters much less.
The model size problem
Speech models improve with size, up to a point, and the largest ones are too big or too slow for many devices.
- Smaller models
- Fit in modest memory and run quickly, even on phones. More errors on accents, noise, rare words and less common languages.
- Larger models
- Better accuracy and language coverage. Need much more memory and compute; slow or impossible on low-end devices.
- Quantization
- Stores model weights with less precision to shrink size and speed up inference, usually with some accuracy cost that varies by model and language.
- Distilled models
- Smaller models trained to imitate larger ones. They narrow the gap but rarely close it across every language.
The practical effect is that on-device quality varies by device. Two people using the same app on different phones may get different speed and, if the app picks a model by hardware, different accuracy. A cloud service runs the same model for everyone. Variants built for speed, such as Whisper large-v3-turbo, show how much engineering goes into the trade between size and accuracy.
Battery, heat and long recordings
Running a neural network continuously is demanding. Transcribing a two-hour recording locally can keep a phone's processor busy for a long time, draining battery and heating the device, and operating systems may slow down a hot phone, stretching the job further. Laptops cope better, especially with a capable GPU, but fans, battery life and other work on the machine still suffer.
Short dictation is cheap by comparison. That asymmetry is why on-device speech has spread fastest for short, interactive tasks and more slowly for bulk transcription of long files.
Local, cloud or hybrid: choosing by situation
- Dictation, voice commands and live captions: on-device, where available for your language.
- Highly confidential audio where no third party may receive it: fully local recognition, accepting lower accuracy or slower processing, or a professional transcriber under agreement.
- Long recordings, many languages or noisy audio where accuracy matters: cloud or hybrid processing with a provider whose retention terms you have checked.
- Weak devices, such as older phones or school laptops: cloud recognition, so results don't depend on hardware.
- Working offline in the field: on-device for notes, then re-transcribe important recordings later if accuracy matters.
A freelance reporter records an interview on her phone. During the conversation she uses on-device live transcription to jot timestamps of key moments. Back home, she wants a clean transcript to quote from, so she runs the full recording through a cloud transcription service with short retention, then checks every quote against the audio before filing.
In that example, each approach did what it is good at. The local transcription was instant and private but rough; the cloud pass was more accurate for the long file. The quotes still needed checking against the recording either way.
Testing a speech tool before you rely on it
- Find out what runs where: is recognition itself local, or only audio preparation? Check the documentation or privacy policy.
- Test on your own typical recordings, not demo audio, including accents, jargon and background noise.
- Try a long file on the device you will actually use and watch time, battery and heat.
- Check language support for every language you need; local options often cover fewer.
- For cloud processing, read the retention period and whether audio is used for training.
- Compare a few minutes of output against a careful manual transcript to judge accuracy.
Trade-offs and limits of local speech models
- Accuracy usually trails larger cloud models, most noticeably in noisy audio, strong accents and less common languages.
- Results depend on hardware, so performance is hard to predict across a team or audience.
- Long files can be slow and drain battery on phones.
- Model downloads can be large, and updating them is up to the app or the user.
- Local does not mean private if the transcript is later synced, shared or backed up to the cloud.
How mydubly splits the work
mydubly uses a hybrid design and says so plainly. Your browser decodes the audio locally to 16 kHz mono, splits it into chunks of about 30 seconds at quiet moments, and compresses each chunk to Opus at about 32 kb/s, roughly 120 KB per 30 seconds. Only those audio chunks are sent over HTTPS; the video file itself never leaves your device, which also keeps uploads small, as explained in reducing upload size for video translation.
Recognition itself does not run on your device. The chunks are transcribed on mydubly's servers with Whisper, so results don't depend on how powerful your phone or laptop is. Audio and results are deleted within 30 minutes of completion, unfinished jobs expire after 24 hours, and the privacy policy says user data is not used for training. If your situation requires that no audio leaves the device at all, mydubly is not the right tool; a fully local recognizer is. More detail on the privacy side is on the private video translation page. For the reporter's 90-minute interview, a transcript and subtitles cost 90 credits (9¢).
Next step
Decide first whether your recordings may leave your device at all. If they may not, use fully local recognition and accept its limits; if short-lived server processing is acceptable, a hybrid tool gives more consistent results on long files. To try the hybrid approach on a recording, use the audio to text tool or the voice memo to text workflow.
Frequently asked questions
Can I run Whisper on my own computer?
Yes. Whisper was released as an open model, and it and community ports can run on many laptops and desktops, with smaller sizes running on some phones. Larger versions need more memory and are much faster with a capable GPU. Setup ranges from command-line tools to desktop apps, so check each project's documentation and hardware requirements.
Is on-device speech recognition less accurate than cloud?
Often, but not always. Accuracy depends on the model, the language and the audio more than on location. A large model running locally on a powerful computer can match a cloud service using the same model. On phones, smaller models are the norm, and the gap is most visible with noise, accents and less common languages.
Does offline dictation on my phone send any audio to the cloud?
It depends on the operating system, the language and your settings. Some phones process dictation fully on-device when the language model is downloaded, and fall back to server processing otherwise or for certain features. Check your device's privacy settings and the manufacturer's current documentation for the specifics.
Why don't all transcription tools just run locally?
Because consistent quality on long files is hard to guarantee across devices. A tool used on old phones and new laptops alike would give very different speed and accuracy if recognition ran locally. Server processing gives everyone the same model, at the cost of sending audio, which is why retention and training policies matter.