Speech recognition & transcription

OpenAI Whisper: how the open speech recognition model works

OpenAI Whisper is a family of speech recognition models released in September 2022 with openly published weights and code. Trained on a very large and varied collection of audio paired with transcripts from the web, a single Whisper model can transcribe speech in many languages, translate speech into English, identify the spoken language and produce timestamps. It is used directly by developers and as the recognition engine inside other products, including mydubly, whose transcription runs on Whisper.

6 min read · Updated

Whisper in a nutshell

Whisper is a transformer encoder-decoder: an encoder reads a log-Mel spectrogram of up to 30 seconds of audio, and a decoder writes text tokens conditioned on it. Architecturally it was deliberately conventional. The novelty lay in the data and the training setup, described in OpenAI's paper "Robust Speech Recognition via Large-Scale Weak Supervision". For the signal-processing side, read how speech to text works first.

The design goal was robustness: a model that works out of the box on audio it was never tuned for, rather than one that scores well on a single benchmark and then struggles with real recordings.

Trained on the open web: weak supervision

Most earlier recognizers were trained on carefully transcribed datasets of a few thousand hours, often read speech recorded in quiet rooms. Whisper's original training set was about 680,000 hours of audio paired with transcripts found on the internet. A substantial share was in languages other than English, and part of it paired non-English speech with English text, which is what taught the translation task.

It is called weak supervision because those transcripts were not produced or checked for the purpose of training. Some were inaccurate, some only loosely matched the audio, and some were themselves machine-generated. OpenAI filtered out transcripts that looked like the output of other recognition systems, but the data stayed noisy. The trade was precision for quantity and diversity: the model heard an enormous range of accents, microphones, rooms and recording qualities.

A side effect of learning from real transcripts is that Whisper absorbed their conventions, such as punctuation, capitalization and numerals, along with some of their habits, like stock phrases common in subtitle files. That is one root of the hallucinations Whisper sometimes produces during silence or music.

One model, several tasks

Instead of training separate models, Whisper encodes the task in special tokens at the start of the decoder's output. The decoder first predicts a language token, then receives a task token, then writes text with or without timestamp tokens.

Language identification
The model predicts which language is spoken by choosing a language token; this is how automatic language detection works
Transcription
Writes the speech as text in the language that was spoken
Translation to English
Writes an English translation of non-English speech; English is the only target Whisper was trained to translate into
Timestamps
Optional tokens marking where segments start and end inside the 30-second window, in 20 ms steps
No-speech detection
An estimate of how likely a window is to contain no speech, which implementations use to skip silent stretches
One Spanish clip, two tasks

Given audio of someone saying "Gracias por venir, empezamos en cinco minutos", the transcribe task returns that Spanish sentence, while the translate task returns something like "Thanks for coming, we'll start in five minutes."

Whisper model sizes

The original release came in five sizes, with English-only versions of the four smaller ones, which tend to do slightly better on English than the multilingual model of the same size.

tiny
About 39 million parameters; very fast, noticeably less accurate
base
About 74 million parameters
small
About 244 million parameters
medium
About 769 million parameters
large
About 1.55 billion parameters; later revised as large-v2 and large-v3
turbo (large-v3-turbo)
About 809 million parameters; large-v3 with a much smaller decoder, released in 2024

Larger models are more accurate, especially on less common languages and difficult audio, but need more memory and compute. The large-v3 revision, released in late 2023, moved to a 128-band spectrogram and was trained on a much larger dataset that included audio labeled by an earlier Whisper model. The turbo variant is covered in its own article, Whisper large-v3-turbo. Parameter counts are as published by OpenAI; the official repository is the place to check current details.

Open weights and the ecosystem around them

OpenAI released Whisper's code and model weights under the MIT license, so anyone can download and run the models, including in commercial products. That openness produced a large ecosystem: the reference Python package, support in Hugging Face Transformers, CPU-friendly ports such as whisper.cpp, optimized runtimes such as faster-whisper, and many fine-tuned variants for particular languages or domains. OpenAI also offers hosted transcription through its API.

For users, the practical consequence is that many transcription apps share the same underlying model. The differences between them come from which size and version they deploy, how they prepare and split the audio, how they deal with silence and how they present the output.

What Whisper does well

  • Handles a wide range of accents, recording conditions and speaking styles without per-user training.
  • Covers close to 100 languages in one model, with strong results in widely spoken ones.
  • Produces punctuated, capitalized, readable text with segment timestamps.
  • Detects the spoken language automatically.
  • Runs almost anywhere: the smaller sizes run on ordinary laptops.

Weaknesses and risks to know about

Whisper is not a streaming model. It processes fixed 30-second windows, so live captioning requires extra engineering and long files have to be split; see transcribing long audio files. Its decoder can hallucinate: in silence, music or noise it may produce plausible sentences, repeated lines or subtitle-style phrases that nobody said. Accuracy also varies a great deal by language, broadly following how much training audio each language had, which multilingual speech recognition explores further.

It also leaves things out. Whisper does not label speakers. It tends to drop filler words and false starts, producing something closer to clean verbatim than true verbatim. Its translation feature only goes into English. And customizing vocabulary is limited to supplying a text prompt or fine-tuning the model yourself.

How mydubly uses Whisper

mydubly's transcription runs on Whisper, and the default deployment uses Whisper large-v3-turbo. Your browser extracts the audio from the file, cuts it at pauses into windows of about 30 seconds, which matches Whisper's native input length, and uploads only the audio. Whisper returns timestamped segments for each window, and the spoken language is detected from the audio, so you only pick the output language: the spoken language itself for an untranslated transcript, or another language for a translation.

Translation is a separate step. A neural machine translation engine translates the recognized segments, which is how mydubly can produce output in any of its 21 languages, whereas Whisper's own translate task only produces English. In video to text you receive the transcript in the spoken language as a timestamped file plus SRT and VTT subtitles, and the same transcript is the basis for translated subtitles and dubbing in the video translator.

Testing Whisper output on your files

  1. Choose two or three representative recordings: one clean, one with background noise, one full of names or jargon.
  2. Open each in video to text and download the transcript.
  3. Check names, numbers and technical terms first, then listen for missing passages and invented lines in quiet stretches.
  4. If you need a number rather than an impression, compute the word error rate against a corrected copy.

Each test minute costs 1 credit, so three five-minute samples cost 15 credits. New accounts created with Google start with 100 free credits, which covers a test like this several times over.

Frequently asked questions

Is Whisper free to use?

The model weights and code are free under the MIT license, so you can run Whisper on your own hardware without licensing fees; you pay for the compute instead. Hosted services built on Whisper, including OpenAI's API and apps like mydubly, charge for processing.

Can Whisper translate between any two languages?

No. Whisper's built-in translation goes from supported languages into English only. Translating into other languages needs a separate machine translation step after transcription, which is how tools that offer many target languages work.

Does Whisper tell speakers apart?

No. Whisper outputs text and timestamps without speaker labels. Attributing lines to speakers requires a separate diarization model, which some tools add on top. mydubly transcripts do not label speakers; speaker diarization explains how that problem is usually approached.

How many languages does Whisper support?

Whisper's tokenizer includes close to 100 languages, but quality varies widely. Languages with plenty of training audio, such as English, Spanish or German, transcribe well, while low-resource languages can be much weaker. Products often offer only the subset where quality is acceptable.

Which Whisper model is most accurate?

Among OpenAI's releases, the large models are the most accurate overall, and large-v3 is the latest full-size version at the time of writing. The turbo variant gives up a small amount of accuracy in exchange for much faster decoding. For a particular language and type of audio, test the candidates on your own recordings.