Speech recognition & transcription

How One Speech Model Transcribes Many Languages, and Why Quality Varies

Multilingual speech recognition uses one neural network, trained on audio in many languages at once, to transcribe all of them. The model shares a single vocabulary and a single set of weights, and is told or decides which language to write. That design makes broad coverage possible, but accuracy is not uniform: languages with lots of training audio come out far cleaner than languages with little.

6 min read · Updated

From one model per language to one model for many

Classic speech recognition systems were built language by language. Each needed its own pronunciation dictionary written by linguists, its own acoustic model trained on carefully transcribed recordings, and its own language model of likely word sequences. Adding a language meant months of work and a fresh data collection effort, so commercial coverage concentrated on large markets.

End-to-end neural models changed the economics. A single encoder-decoder transformer can learn the mapping from audio to text for dozens of languages simultaneously, as long as it sees enough paired examples. OpenAI's Whisper, released with open weights in 2022, is a widely used example: one checkpoint covers roughly a hundred languages, and the same network handles transcription, translation into English, language identification and timestamps. For the general background, see what OpenAI Whisper is.

A shared vocabulary across scripts

A multilingual model has to write Latin, Cyrillic, Arabic, Devanagari, Hangul, Chinese characters and more from one output layer. Whisper does this with a byte-level byte-pair-encoding tokenizer: text is broken into frequently occurring pieces, and anything not covered falls back to raw bytes. Every script can therefore be expressed, but not equally efficiently.

Common English words often map to a single token. Words in scripts that were less frequent when the tokenizer was built may be split into several tokens, sometimes down to individual bytes. More tokens per word means a longer sequence for the decoder to produce, more chances to slip, and fewer words fitting into the model's fixed output window. This is one quiet reason why quality can differ between languages even when audio quality is identical.

How the model knows which language to write

Whisper's decoder is steered by special tokens. After a start-of-transcript token comes a language token, such as one for English or Japanese, followed by a task token for transcription or translation. The language token can be supplied by the application, or the model can predict it from the audio before it starts writing. That prediction step is how automatic detection works, and its failure modes are covered in the article on spoken language identification.

Because all languages share the same weights, knowledge transfers between them. Sounds, speaking patterns and even vocabulary learned from Spanish help with Portuguese and Italian; Hindi benefits from related Indo-Aryan data. The flip side is interference: closely related languages can bleed into each other, and the model can drift into the wrong language mid-file.

Why accuracy varies so much by language

Four factors explain most of the gap between a model's strongest and weakest languages.

Training data volume
The single biggest factor. The Whisper paper shows error rates falling steadily as the hours of training audio for a language rise.
Writing system
Languages without spaces between words, such as Chinese and Japanese, are measured with character error rate rather than word error rate, so their scores are not directly comparable to English.
Morphology
In agglutinative languages such as Finnish and Turkish, one word carries what English spreads across several, so a single wrong suffix counts as a whole wrong word.
Spoken versus written gap
Where everyday speech differs from the standard written form, as with Arabic dialects and Modern Standard Arabic, the model must both recognize and normalize, and spelling can vary.

The word error rate explainer covers how these scores are computed and why a number published for one language tells you little about another.

Low-resource languages and what transfer can and cannot do

A low-resource language is one with little transcribed audio available, regardless of how many people speak it. Multilingual training is what makes such languages usable at all: the model borrows acoustic knowledge from related languages and general knowledge of speech from everything else. A language with a few thousand training hours in a multilingual model can outperform what a dedicated model trained on those hours alone would achieve.

Transfer has a ceiling, though. It cannot invent vocabulary the model never saw written, and it cannot teach correct spelling conventions for a language whose online text is sparse or inconsistent. For low-resource languages, expect usable drafts that need real editing rather than near-final text.

Where multilingual models clearly help

  • One tool and one workflow for every language you publish in, instead of separate vendors per market.
  • Reasonable quality on mid-sized languages that traditional systems never covered.
  • Robustness to accented speech, because the model has heard many languages' sound systems.
  • Built-in language identification, so batches of mixed-language files need no manual sorting.
  • Speech-to-English translation in the same model, which Whisper offers as a separate task.

Trade-offs and limitations of the one-model approach

  • Quality is uneven, and the language you need may sit in the weaker half of the range.
  • Code-switching, where speakers mix languages in a sentence, is handled inconsistently; the model tends to commit to one language per stretch of audio.
  • Closely related languages can be confused, especially on short clips.
  • Script choices may not match your expectations, for example simplified versus traditional Chinese characters.
  • You cannot easily add vocabulary for one language without retraining or fine-tuning the whole model.

Why mydubly supports these 21 languages

mydubly's transcription runs on Whisper, with large-v3-turbo as the default model, and its default voice engine is Chatterbox Multilingual, an open-source text-to-speech model from Resemble AI. Whisper recognizes many more languages than mydubly lists, but a full mydubly job can end with a translated voice track, so the product offers the 21 languages that both its speech recognition and its voice output support: English, Arabic, Chinese, Czech, Dutch, Finnish, French, German, Hebrew, Hindi, Italian, Japanese, Korean, Norwegian, Polish, Portuguese, Russian, Spanish, Swedish, Turkish and Vietnamese.

The spoken language is detected automatically, so you never select a source language. Written Chinese translation output is Simplified Chinese, Arabic is Modern Standard Arabic and Norwegian is Bokmål. Each language has its own page, such as Japanese video to text or Arabic video to text.

Test a language before committing a large project

Published scores rarely match your material, so a short trial on your own recordings is the most reliable guide.

  1. Pick a representative three-to-five-minute clip for each language, including your typical speakers, pace and background.
  2. Transcribe each clip and correct it by hand to create a reference.
  3. Note error types, not just counts: names, numbers, technical terms, wrong language or wrong script.
  4. Decide per language whether output is publishable after a skim, needs a full proofread, or needs human transcription.
  5. Budget proofreading time per language accordingly before scaling up.
Example: a three-language course

Suppose a training team publishes lessons in English, Polish and Vietnamese. They transcribe a 5-minute sample in each language, which costs 15 credits in total (1.5 cents). English needs two fixes, Polish a handful of case-ending corrections, and Vietnamese several tone-mark and name fixes. They plan a light skim for English and a full proofread for the other two before translating.

Next step

Run your own trial on video to text with a few minutes of audio in each language you work in. New accounts that sign up with Google get 100 free credits, enough to transcribe up to 100 minutes of test material across several languages.

Frequently asked questions

How many languages can a multilingual speech model handle?

It depends on the model. Whisper covers roughly a hundred languages, though quality across that list ranges from excellent to rough. Check the model's official documentation for the current list, and test the languages you need on your own audio.

Is a dedicated single-language model more accurate than a multilingual one?

Sometimes, for a language with abundant training data and a model built specifically for it. For most mid-sized and smaller languages, multilingual training helps because the model borrows knowledge from related languages.

Why does mydubly list 21 languages when Whisper recognizes more?

Because a mydubly job can include a translated AI voice, the language list is limited to languages that both its speech model and its voice model support. That keeps every listed language usable for transcripts, subtitles and dubbing.

Can a multilingual model transcribe a video where two languages are spoken?

Partly. Models like Whisper tend to settle on one language for a stretch of audio, so switched-in words or sentences in another language may be mistranscribed or translated. Treat mixed-language files as needing a careful proofread.

Why are Chinese and Japanese measured differently from English?

They are written without spaces between words, so splitting text into words is itself ambiguous. Character error rate, which counts errors per character, gives a fairer measure for those languages.