The pipeline at a glance
- Capture and sample the sound wave, then resample it to the rate the model expects, which is 16 kHz for Whisper.
- Slice the audio into short overlapping frames and measure each frame's energy across Mel-spaced frequency bands, on a log scale.
- Feed the resulting spectrogram to an encoder network that builds a contextual representation of the audio.
- Let a decoder network generate text tokens one at a time, each conditioned on the audio and on the tokens already written.
- Search over likely token sequences, then turn tokens back into text with punctuation, casing and timestamps.
The sections below use Whisper as the running example because its design is public; systems built on CTC or transducer decoders differ mainly in the back half.
Sampling: from air pressure to numbers
A microphone turns changes in air pressure into a voltage, and an analog-to-digital converter measures that voltage thousands of times per second. Each measurement is a sample. Music is usually recorded at 44.1 or 48 kHz; speech recognition models typically work at 16 kHz. By the Nyquist theorem, a 16 kHz sample rate captures frequencies up to 8 kHz, which covers the range carrying most of the information that distinguishes speech sounds, including the hiss of consonants like "s" and "f".
Recognition models also use a single channel. Stereo is mixed down to mono, and each sample is stored as a number, commonly a 16-bit integer or a 32-bit float between minus one and one.
One second of 16 kHz, 16-bit mono audio is 16,000 samples × 2 bytes = 32,000 bytes. A 30-second window is 480,000 samples, about 960 KB. A two-hour recording is 115.2 million samples, roughly 230 MB, before any compression.
Log-Mel spectrograms: what the model actually sees
Raw samples are a poor input for a neural network: the sequence is very long, and the same word looks completely different depending on pitch and volume. So the audio is converted into a spectrogram. The signal is cut into short frames (Whisper uses 25-millisecond windows advancing 10 milliseconds at a time) and a Fourier transform measures how much energy each frame contains at each frequency.
Those frequencies are then grouped into Mel bands. The Mel scale spaces bands roughly the way human hearing does: narrow at low frequencies, where we notice small pitch differences, and wide at high frequencies. Whisper's original models use 80 Mel bands; large-v3 moved to 128. Finally the energies are put on a logarithmic scale, which compresses the huge range between a whisper and a shout into values a network handles comfortably, and then normalized.
The result is a two-dimensional array with time along one axis and frequency bands along the other. A 30-second window at a 10 ms hop becomes 3,000 frames. Vowels appear as horizontal bands of energy called formants, plosives like "p" and "t" as short vertical bursts, and fricatives as fuzzy high-frequency noise.
The encoder-decoder transformer
Whisper's encoder begins with two small convolutional layers that detect local patterns and halve the time resolution, turning 3,000 frames into 1,500 positions. Positional information is added so the model knows the order of frames, and then a stack of transformer layers processes the sequence. Self-attention lets every position look at every other position in the window, so the representation of an ambiguous sound is informed by what comes before and after it.
The decoder is a second transformer stack. It reads the tokens generated so far and, through cross-attention, consults the encoder's output to decide which part of the audio matters for the next token. The encoder runs once per 30-second window, while the decoder runs once for every token it produces. That asymmetry is why Whisper large-v3-turbo could be made much faster by shrinking only the decoder.
Tokens and the decoding loop
The decoder does not output letters or whole words but tokens from a fixed vocabulary of word pieces built with byte-pair encoding. A common word like "the" is a single token; a rare surname may be split into several pieces. This lets a finite vocabulary spell any word in any script.
Whisper also uses special tokens as instructions. An output sequence starts with a start-of-transcript token, then a language token, then a task token (transcribe, or translate into English), and optionally a token that switches timestamps off. The model then predicts text tokens until it emits an end-of-text token. Because language identification is simply predicting the language token, the same model can detect the spoken language before transcribing it.
At each step the decoder assigns a probability to every token in the vocabulary. Greedy decoding just takes the most likely one. Beam search keeps several partial transcripts, called the beam, extends each of them and retains the highest-scoring sequences overall. That avoids committing early to a word that looks likely in isolation but leads to a worse sentence. Whisper's reference implementation adds a fallback: if the output looks unreliable, judged by a low average log probability or highly repetitive text, it decodes the window again with sampling at a higher temperature.
Timestamps, punctuation and casing
Older pipelines produced lowercase words without punctuation and relied on separate models to restore punctuation, capitalize sentences and convert spoken forms like "twenty five dollars" into "$25", a step called inverse text normalization. Whisper learned all of this from its training transcripts, so its decoder writes punctuation, capitals and digits directly.
Timestamps work the same way. Whisper has timestamp tokens that represent positions within the window in 20-millisecond steps. With timestamps on, the decoder writes a start-time token, the segment's text and an end-time token. Word-level timings, where a tool offers them, are usually derived afterwards by aligning tokens with the audio using the model's attention patterns.
A classic pipeline might emit "so the meeting is on march third at two thirty pm in room four b". An end-to-end model trained on formatted text tends to write "So the meeting is on March 3rd at 2:30 p.m. in Room 4B." The exact formatting choices vary between models and even between runs.
What this design gets right
- One network learns acoustics, pronunciation and language together, so there is no pronunciation dictionary to maintain.
- Attention across a full 30-second window gives the model sentence-level context for ambiguous sounds.
- Training on varied real-world audio makes the model more tolerant of accents and moderate noise than models trained on clean read speech.
- Punctuation, casing, numerals and timestamps arrive in one pass, producing readable transcripts and usable subtitle timing.
- The same model can identify the language and then transcribe it.
Limits: where speech to text goes wrong
- The decoder is a language model, so when the audio is unclear it can write fluent text that was never said, particularly during silence or music. This is covered in Whisper hallucinations.
- The 30-second window means long recordings must be chunked, and words near a cut can be lost or duplicated if the cut is badly placed.
- Tokens for rare names are improbable by definition, so the model often substitutes a similar-sounding common word.
- Beam search improves accuracy but costs compute, so long transcripts take real processing time.
- Noise and reverberation smear the spectrogram and hide the cues the encoder relies on.
How mydubly runs speech to text
In mydubly, the first two stages of this pipeline begin in your browser. A WebAssembly build of ffmpeg decodes the file's audio once into 16 kHz mono 16-bit PCM, exactly the format described above. The audio is then cut into windows of about 30 seconds, with each cut placed at the quietest 50 ms frame within the last six seconds before the 30-second mark, so cuts land in pauses rather than in the middle of a word.
Each window is compressed to Ogg Opus at 32 kb/s, about 120 KB per 30 seconds, and uploaded. On the server, Whisper large-v3-turbo computes the spectrogram, runs the encoder and decoder, and returns text segments with start and end timestamps for each chunk. Those segments become your timestamped transcript and your SRT or VTT files; the timestamped transcript page shows what that output looks like.
Try it on your own audio
The quickest way to build intuition is to compare a clean file with a difficult one. Upload a short recording to audio to text and read the result against the audio. A transcript costs 1 credit per minute with a 5-credit minimum per file, so a ten-minute test is 10 credits, one cent. For recording advice that improves the input, see how to improve transcription accuracy.
Frequently asked questions
Why do speech recognition models use 16 kHz audio?
Sampling at 16 kHz captures frequencies up to 8 kHz, which includes most of the acoustic detail that distinguishes speech sounds. Higher rates mostly add information that matters for music rather than words, while making inputs larger. That's why recordings are resampled to 16 kHz before recognition, whatever rate they were recorded at.
What is a log-Mel spectrogram in simple terms?
It is a picture of sound: time runs left to right, frequency bands run bottom to top, and brightness shows how much energy sits in each band at each moment. The bands are spaced like human hearing and the energies are on a log scale. Speech models read this picture rather than the raw waveform.
Does beam search make transcription more accurate?
Usually somewhat. Keeping several candidate transcripts lets the decoder recover from a choice that was locally likely but globally wrong, which helps with ambiguous words. It costs extra computation, and on very clear audio greedy decoding often produces the same result.
Where do punctuation and capital letters come from?
In end-to-end models like Whisper, the decoder predicts punctuation and capitals as ordinary tokens because its training transcripts contained them. Older systems produced bare words and added punctuation with a separate model. Either way punctuation is a prediction, so check it wherever meaning depends on it.
How are segment timestamps produced?
Whisper predicts special timestamp tokens, in 20-millisecond steps, before and after each segment of text. Those times are relative to the 30-second window, so a system processing long audio adds each window's offset to place segments on the full timeline. That offsetting is covered in subtitle timestamp alignment.