The two architectures side by side
- Building blocks
- Hybrid: acoustic model, pronunciation lexicon, n-gram language model and a decoding graph. End-to-end: one encoder-decoder network
- Training data
- Hybrid: transcribed audio aligned to phonetic states, plus large text corpora. Whisper: hundreds of thousands of hours of audio paired with transcripts
- Output
- Hybrid: lowercase words, usually without punctuation until post-processing. Whisper: punctuated, cased text with numerals
- Adding a new word
- Hybrid: add it to the lexicon and language model. Whisper: prompt it or fine-tune the model
- Real-time use
- Hybrid: commonly built for streaming. Whisper: built around 30-second windows
- Typical failure
- Hybrid: misrecognized or out-of-vocabulary words. Whisper: occasionally fluent text that was never spoken
Inside a hybrid HMM/DNN pipeline
A classic system starts with acoustic features, historically MFCCs or filterbank energies computed every 10 milliseconds. An acoustic model, which in the hybrid era is a deep neural network, estimates for every frame how likely it is to belong to each of thousands of context-dependent phonetic states. Hidden Markov models describe how those states follow one another to form a phoneme and how phonemes chain into words.
The pronunciation lexicon lists every word the system can recognize with one or more phoneme sequences, so "tomato" might have two entries. A language model, typically an n-gram model trained on text from the target domain, scores how likely each word sequence is. These parts are compiled into one large search graph, often with weighted finite-state transducers as in the open-source Kaldi toolkit, and a decoder searches that graph for the best path through the audio.
Every component is a separate engineering artifact with its own training data, tools and failure modes. That is both the weakness and the strength of the approach.
Inside an end-to-end model like Whisper
Whisper has no lexicon, no separate language model and no phonetic states. Its encoder turns a log-Mel spectrogram into representations, and its decoder writes word-piece tokens one by one while attending to the audio. Pronunciation, spelling, grammar, punctuation and even language identification are learned implicitly from paired audio and text. The technical walk-through of speech to text goes through each stage.
Because everything is learned jointly, there is nothing to hand-tune, but also nothing to reach into. You cannot open Whisper and add a pronunciation for a new product name the way you would add a lexicon entry.
Advantages of the end-to-end approach
- Robustness: trained on varied web audio, Whisper copes with accents, microphones and moderate noise that would need separate adaptation in a hybrid system.
- Readable output: punctuation, casing and numerals are built in, with no separate restoration or normalization models.
- Many languages in one model: a hybrid system needs a lexicon and language model per language, often built with linguists.
- Low maintenance: one model file replaces a toolchain of lexicons, graphs and feature pipelines.
- Open weights: the models can be downloaded, run and fine-tuned by anyone.
Trade-offs: where traditional pipelines still make sense
Customization is the clearest case. If a call center must recognize thousands of product codes, or a clinic needs drug names spelled exactly, a hybrid system lets engineers add words to the lexicon and boost them in the language model without retraining the acoustic model. Constrained grammars, where only certain phrases are valid, as in phone menus, are natural in a hybrid decoder and awkward in Whisper.
Streaming is the second. Hybrid and transducer systems can emit words a fraction of a second after they are spoken. Whisper waits for a window of audio and then decodes it; streaming wrappers exist, but they add latency and complexity. For live captions or voice commands, purpose-built streaming models remain the usual choice.
Footprint and traceability are the third. Small hybrid models can run on modest embedded hardware, and their errors can be traced to a specific component. When a hybrid system fails, an engineer can ask whether the word was missing from the lexicon or scored low by the language model. When Whisper fails, the diagnosis is often just that the model decided otherwise.
Hallucination risk versus recognizable errors
The two families fail differently. A hybrid system can only output words in its lexicon, and its decoder must account for the audio frame by frame, so when it fails it usually produces visibly wrong words, the classic example being "recognize speech" heard as "wreck a nice beach". Those errors are annoying but easy to spot.
Whisper's decoder is a powerful language model, so its errors tend to be fluent. Given silence, music or heavy noise, it may write a plausible sentence, repeat a line several times or produce subtitle-style phrases nobody said. Fluent errors are harder to catch when proofreading, which is why mitigation matters: voice activity detection to skip non-speech, cutting audio at pauses and checking quiet stretches. Whisper hallucinations covers causes and fixes in detail.
Suppose a podcast ends with 20 seconds of music. A hybrid recognizer might output a few scattered short words such as "and the uh in". A Whisper-style model might instead output a tidy line like "Thanks for watching!" that appears nowhere in the audio. The first is obviously junk; the second looks like real speech.
Customizing vocabulary in each world
- Hybrid: add the term and its pronunciation to the lexicon, add example sentences to the language model's training text, rebuild the decoding graph and test.
- Whisper, light touch: pass a prompt containing the correct spellings of names and terms, which nudges the decoder toward them without guaranteeing them.
- Whisper, heavy touch: fine-tune the model on transcribed audio from your domain, which requires data, compute and an evaluation set.
- Either system: correct the transcript afterwards with find-and-replace for recurring terms, often the most practical option for a handful of names.
Choosing for a real project
For recorded content such as interviews, lectures, podcasts and videos, where accuracy across varied audio matters more than latency, an end-to-end model is usually the pragmatic choice. For voice interfaces, live captioning, tightly constrained vocabularies or very small devices, a streaming or hybrid system may still fit better. Many organizations run both: a streaming model for live use and a Whisper-class model for the archive.
Why mydubly uses Whisper
mydubly processes recorded files rather than live streams, so Whisper's window-based design fits the job. Its transcription runs on Whisper, with large-v3-turbo as the default model. To reduce the boundary and silence problems described above, the browser cuts each window of about 30 seconds at the quietest 50 ms frame within the last six seconds before the cut point, so chunks start and end in pauses rather than mid-word.
What mydubly doesn't offer is the hybrid-style toolkit: there are no custom vocabularies, no word boosting and no speaker labels. For recurring names, the practical approach is to check and correct them in the transcript. Transcripts cost 1 credit per minute, and video to text handles the same job for video files.
Next step: judge the output, not the architecture
Architecture debates settle quickly once you see results on your own recordings. Run a few representative files through audio to text, including one heavy with jargon and one with quiet or musical passages, and note where the errors land. If you want to put a number on the result, word error rate explains how to calculate it.
Frequently asked questions
Is Whisper always more accurate than a traditional recognizer?
Not always. On varied real-world recordings it is usually more robust out of the box, but a hybrid system carefully tuned to a narrow domain, with vocabulary and language model built for it, can do better on that domain's jargon. The only reliable comparison is on your own audio.
Can I add custom words to Whisper?
Not in the lexicon sense, because Whisper has no lexicon. You can pass a text prompt with the correct spellings, which biases the decoder toward them, or fine-tune the model on domain audio. mydubly does not offer custom vocabularies, so recurring terms are best fixed in the transcript.
Why isn't Whisper used for live captions more often?
It was designed to process 30-second windows, so producing text while someone speaks requires buffering, repeatedly decoding overlapping windows and logic to stabilize the output. That adds latency and complexity. Purpose-built streaming models emit words sooner and are the usual choice for live use.
What is Kaldi?
Kaldi is an open-source toolkit for building speech recognition systems that was widely used in research and industry during the hybrid HMM/DNN era. It provides tools for feature extraction, acoustic model training, lexicons, language models and finite-state decoding graphs, and it typifies the component-based approach that end-to-end models replaced in many applications.
Do traditional systems hallucinate?
Not in the way Whisper can. A hybrid system can only output words in its lexicon and must account for the audio frame by frame, so failures usually look like wrong or garbled words rather than invented fluent sentences. They can still produce short spurious words in noise.