Speech recognition & transcription

Whisper large-v3-turbo: a faster Whisper with a four-layer decoder

Whisper large-v3-turbo is a version of OpenAI's Whisper large-v3 speech recognition model with its decoder cut from 32 transformer layers to 4 and then fine-tuned. The encoder that listens to the audio is unchanged, so the model keeps most of large-v3's accuracy while generating text much faster. OpenAI released it in 2024, and it is the default speech model behind mydubly's transcripts.

6 min read · Updated

What large-v3-turbo actually is

Whisper models have two halves. The encoder turns a 30-second log-Mel spectrogram into a sequence of acoustic representations; the decoder writes the transcript token by token while attending to those representations. In Whisper large-v3 both halves have 32 transformer layers. Turbo keeps all 32 encoder layers and reduces the decoder to 4, bringing the parameter count from about 1.55 billion down to about 809 million.

Simply deleting layers would damage the model, so OpenAI fine-tuned the pruned model afterwards on transcription data. According to OpenAI's documentation, turbo was not trained for the translation task, which matters if you planned to use Whisper's speech-to-English translation; more on that below. The tokenizer, the spectrogram format (128 Mel bands, as in large-v3), the language list and the timestamp tokens are unchanged, so turbo is a drop-in replacement in most code that already runs large-v3.

Where turbo sits in the Whisper lineage

large (2022)
The original largest Whisper model, about 1.55 billion parameters, using 80 Mel bands
large-v2 (late 2022)
Same architecture, trained for longer with added regularization
large-v3 (late 2023)
128 Mel bands, a new Cantonese language token and a larger training set that included audio pseudo-labeled by large-v2
large-v3-turbo (2024)
The large-v3 encoder with a 4-layer decoder, about 809 million parameters, fine-tuned after pruning

The broader background on the family, including how it was trained and what tasks it handles, is in what is OpenAI Whisper. This article focuses on what changed in turbo and what that means in practice.

Why a smaller decoder makes it so much faster

The two halves of the model run a different number of times. The encoder processes each 30-second window once, in parallel across all positions. The decoder is autoregressive: it runs once per output token, and each run depends on the one before, so those steps cannot be parallelized. A 30-second window of fast speech can easily need around a hundred tokens, including timestamp tokens, which means the decoder runs about a hundred times for every encoder pass.

Rough arithmetic

Suppose a window needs 100 decoder steps. With large-v3 that is 100 passes through 32 decoder layers, or 3,200 layer evaluations, after one pass through 32 encoder layers. With turbo it is 100 passes through 4 layers, or 400 layer evaluations, after the same encoder pass. The encoder cost stays the same, but the sequential part of the work shrinks by a factor of eight.

Real-world speedups are smaller than the layer count suggests, because the encoder, memory transfers and beam search overhead don't shrink. OpenAI's README describes turbo as several times faster than large; the exact factor depends on the hardware, batch size, runtime and decoding settings. Faster decoding also means a lower cost per minute for anyone running the model at scale.

Keeping a strong encoder and trimming the decoder had been explored earlier by Hugging Face's Distil-Whisper project, which used knowledge distillation to train much smaller decoders. The two efforts differ in method and language coverage, so check each model card before choosing.

Trade-offs: accuracy, translation and harder languages

Turbo is slightly less accurate than large-v3. The encoder hears the audio just as well, but a 4-layer decoder has less capacity for the language-modeling side of the job: choosing between similar-sounding words, spelling rare names and keeping long sentences coherent. OpenAI's published per-language comparisons show the gap is small for many widely spoken languages and larger for some others, so the fair summary is close to large-v3, not identical.

Translation is the bigger caveat. OpenAI's README states that turbo was not trained for translation and points users to the medium or large models for that task. The restriction doesn't affect transcription, and it doesn't affect pipelines that translate the transcript text with a separate engine afterwards.

Turbo also inherits large-v3's general weaknesses unchanged. It can hallucinate in silence or music, it works in 30-second windows, it does not label speakers, and it is still a large model that needs a capable GPU for fast inference. Some users have reported that different Whisper versions hallucinate differently on the same audio, so if a workflow depends on one model's behavior, test before switching.

When turbo is a good choice

  • Transcribing large volumes of audio where throughput and cost matter as much as the last increment of accuracy.
  • Widely spoken languages with plenty of training data, where the gap to large-v3 is smallest.
  • Pipelines that transcribe first and then translate text with a dedicated translation engine.
  • Interactive tools where people wait for results and every second of latency is noticed.
  • Hardware with limited memory, where the smaller parameter count helps.

When to choose a different Whisper model

Use large-v3 if you need Whisper's built-in translation into English, if your language is one where turbo's published results trail large-v3 noticeably, or if your audio is very hard (strong accents, poor microphones, dense jargon) and you can afford slower decoding. Use the English-only small or medium models when you only transcribe English and must run on modest hardware. If you need live captions with low latency, look at streaming-oriented architectures, since every Whisper model, turbo included, is built around 30-second windows; Whisper vs traditional speech recognition covers that trade-off.

Running large-v3-turbo yourself

The reference openai-whisper Python package exposes the model under the name "turbo". On Hugging Face it is published as openai/whisper-large-v3-turbo and works with the Transformers speech recognition pipeline. Community runtimes such as whisper.cpp and faster-whisper offer converted versions for CPU and GPU inference. Names and options change between releases, so confirm the current identifiers in each project's documentation.

  1. Pick a runtime that matches your hardware: a GPU runtime for throughput, or whisper.cpp for CPU-only machines.
  2. Resample input to 16 kHz mono, which every Whisper runtime expects internally.
  3. Leave the task set to transcribe, and set the language explicitly if you already know it, to avoid misdetection on short clips.
  4. For files longer than 30 seconds, use the runtime's long-form mode or split the audio at pauses yourself.
  5. Compare turbo with large-v3 on a few of your own files before committing to either.

Turbo inside mydubly

mydubly's default deployment transcribes with Whisper large-v3-turbo. The browser decodes your file's audio to 16 kHz mono, cuts it into windows of about 30 seconds at the quietest moment near each boundary, and uploads several chunks in parallel; turbo returns text segments with start and end timestamps for each one. Because translation is handled afterwards by a separate neural machine translation engine working on the transcript, turbo's lack of translation training never comes into play.

Transcripts cost 1 credit per minute, so a 50-minute webinar is 50 credits, or five cents. You get a timestamped transcript plus SRT and VTT subtitles from video to text, and the spoken language is detected automatically from the audio.

Next step: hear turbo on your own audio

Model comparisons only go so far; what matters is how the model handles your speakers, your microphones and your vocabulary. Open a representative recording in video to text or audio to text, read the transcript against the audio, and pay particular attention to names and numbers. If those need fixing, how to proofread an AI transcript describes an efficient review routine.

Frequently asked questions

Is Whisper large-v3-turbo more accurate than large-v3?

No. It is slightly less accurate because its decoder has 4 layers instead of 32. The encoder is identical, so the gap is small for many languages, and the trade buys much faster transcription. For the hardest audio or less common languages, large-v3 may be worth the extra time.

Can Whisper turbo translate audio into English?

OpenAI's documentation says turbo was not trained for translation and recommends the medium or large models for that task, so don't rely on it there. Alternatively, transcribe with turbo and translate the resulting text with a separate translation engine, which is how many multilingual pipelines work.

How much faster is large-v3-turbo?

OpenAI describes it as several times faster than the large model, mostly because the decoder, which runs once per output token, is eight times shallower. The real speedup depends on your hardware, batch size, runtime and decoding settings, so measure it in your own setup.

Does turbo support the same languages as large-v3?

It uses the same tokenizer and language tokens, so the list of languages it can attempt is the same. Accuracy per language is not identical, though. Check OpenAI's per-language comparisons and test the specific languages you care about.

Is large-v3-turbo the same as Distil-Whisper?

No. Both keep a large Whisper encoder and shrink the decoder, but they come from different teams and were trained differently: turbo is OpenAI's pruned and fine-tuned large-v3, while Distil-Whisper, from Hugging Face, uses knowledge distillation. Their model cards describe different language coverage and trade-offs.