What the kb/s number actually means
A bitrate of 128 kb/s means the encoded stream uses 128,000 bits for each second of audio. Divide by eight to get bytes: 16,000 bytes per second, or 0.96 MB per minute. That is the whole budget the encoder has for describing the sound, shared across all channels. A 128 kb/s stereo file gives each channel roughly half; the same 128 kb/s in mono spends it all on one signal.
Two units cause constant confusion. Kilobits per second (kb/s or kbps) describe audio streams. Kilobytes per second (KB/s) describe download and file speeds, and are eight times larger. A 64 kb/s podcast stream therefore needs only about 8 KB per second of network capacity.
Uncompressed audio has a bitrate too, set entirely by its format: 48 kHz, 16-bit stereo PCM runs at 1,536 kb/s. Compression exists to bring that down by a large factor, and the bitrate you choose decides how much the codec has to throw away to get there. If the difference between lossy and lossless coding is new to you, start with what an audio codec is.
Constant, variable and average bitrate
Encoders spend their budget in different ways:
- CBR
- Constant bitrate. Every second gets the same number of bits, whether it is silence or a burst of speech. Predictable size, wasteful on easy passages
- VBR
- Variable bitrate. The encoder aims for a quality level and spends more bits on complex sounds and fewer on simple ones. Usually better quality per megabyte
- ABR
- Average bitrate. Varies moment to moment but keeps to a target average. A compromise between the two
- Nominal bitrate
- The figure a player shows for a VBR file, typically an average, so two files showing the same number can differ in quality
For speech, VBR generally serves well, because pauses and steady vowels need far fewer bits than sharp consonants. CBR still has its place where a system needs a predictable data rate, such as some streaming setups.
What low bitrates do to a voice
As the budget shrinks, a lossy encoder has to make harsher choices. The audible results are distinctive once you know them:
- A dull top end. Many encoders cut high frequencies entirely at low bitrates to save bits for the core of the sound, and the s, f and t sounds lose their crispness.
- Swirling or watery texture. Coarse quantization makes the background and the voice itself warble, sometimes described as sounding underwater.
- Smeared attacks. Sharp onsets such as k, p and t blur into the sounds around them, an artefact known as pre-echo.
- Synthetic highs. Some low-bitrate modes rebuild the top end from the lower frequencies instead of storing it, which can sound hissy or artificial.
The same bitrate does not mean the same quality across codecs. Newer designs such as AAC and especially Opus generally hold speech together at bitrates where older MP3 encoders struggle, which is why voice apps tend to use them. Opus in particular was designed with speech in mind; the Opus audio codec article explains how.
How speech recognition responds
A model such as Whisper is trained on a very large amount of real-world audio, much of it compressed, so it tolerates moderate compression well. Ordinary podcast, phone-camera and video-platform audio is well within what it has learned to handle. Problems grow when compression removes the features that separate one word from another.
The damage is not spread evenly. Vowels survive heavy compression fairly well, while consonants carry much of the high-frequency detail an encoder drops first. Errors therefore concentrate in exactly the places that matter most for a usable transcript: proper names, unfamiliar terms, numbers and short function words. Compression also interacts with other problems. A clean voice at a low bitrate may be fine, while the same bitrate applied to a voice already competing with room noise, covered in speech recognition and background noise, can tip a recording into guesswork.
Why re-encoding at a higher bitrate can't help
A common instinct is to convert a small, poor-sounding file into a high-bitrate MP3 or a WAV before transcribing it. It does not work. Decoding the original reproduces exactly the degraded sound the first encoder left behind, and the new encode, however generous, can only preserve that degraded sound. The file grows several times larger and sounds the same, or very slightly worse if the new encode is lossy.
You can see this for yourself. Open both versions in an audio editor's spectrogram view: the high-frequency region that the first encoder removed is empty in both. The same applies to video soundtracks extracted and re-exported several times, and to recordings passed through messaging apps that recompress them.
The extra encode in mydubly's pipeline
mydubly adds one lossy step of its own, and it is worth being clear about why that is reasonable. In your browser, the audio is decoded once to 16 kHz mono and split into chunks of about 30 seconds, and each chunk is compressed to Opus at about 32 kb/s before it is sent over HTTPS to be transcribed with Whisper. The video file stays on your device.
Thirty-two kilobits per second is a comfortable budget for a single channel of 16 kHz speech with a codec built for voice, and the recognizer only uses content up to 8 kHz in any case. For a clean source the effect on the transcript is generally small. For an already heavily compressed source, the extra encode cannot make it better and adds a little further approximation; the source's own damage remains the dominant factor. Where a browser cannot encode Opus, the chunks are sent as uncompressed WAV instead. The size and upload-time side of this is covered in reducing upload size.
A reporter has a 40-minute interview as a 128 kb/s AAC file from her recorder, and a copy a colleague sent through a messaging app at a far lower bitrate. She cuts the same two-minute section, with several names and figures, from each and transcribes both separately, at 5 credits (0.5¢) each because of the per-file minimum. The recorder copy gets every name right; the compressed copy misses a surname and turns fifteen into fifty. She transcribes the full interview from the recorder file.
Choosing and checking bitrate for speech
- Use the original recording whenever you have it. The first copy off the recorder or camera is almost always the best.
- When exporting from an editor, choose a lossless format such as WAV or FLAC, or a lossy one at a moderate bitrate; around 96 to 128 kb/s for AAC or MP3 is a common, comfortable choice for speech.
- For single-voice recordings, export mono so the whole budget goes to one channel; mono vs stereo for speech explains when that is safe.
- Avoid sharing source files through apps that recompress media. Use file transfer or cloud storage that keeps the original bytes.
- When a file's quality is in doubt, transcribe a short section with names and numbers and compare it with the audio before processing the whole thing.
Limits of what bitrate can tell you
- A high bitrate does not guarantee a good recording. A distant, echoey or noisy voice stays that way at 320 kb/s.
- The displayed bitrate may hide history. A 192 kb/s MP3 made from a 48 kb/s source is still a 48 kb/s recording in all but size.
- Codec matters as much as the number. Compare bitrates only within the same codec and similar settings.
- Very low bitrates can be fine for listening comprehension and still hurt transcripts, because people use context more freely than models do.
- No setting on export recovers clipping, wind or hum; those need fixing at the recording stage.
Next step
Check the codec and bitrate of the recordings you transcribe most often, and keep the originals rather than shared copies. To see how a particular file performs, run a short section through audio to text or, for video, video to text; the MP3 to text page covers the most common compressed format. For measuring the difference between two versions more formally, word error rate explains how.
Frequently asked questions
What bitrate is good enough for transcribing speech?
There is no single cutoff, because it depends on the codec and the recording. As a practical guide, clean speech in AAC or MP3 at roughly 64 kb/s and above, especially in mono, usually transcribes well, and modern codecs such as Opus stay clear at lower rates. Files that have been recompressed several times, or that already sound watery, are where accuracy tends to drop.
Does a 320 kb/s MP3 transcribe better than a 128 kb/s one?
For speech recorded cleanly, usually not by any noticeable amount. At 128 kb/s an MP3 encoder already keeps nearly all of the detail that distinguishes words, and the recognizer resamples to 16 kHz, discarding much of what the extra bits describe. The higher bitrate makes more sense for music or for a master copy you might edit later.
Is variable bitrate better than constant bitrate for voice?
Generally yes, for files. Speech alternates between pauses, steady vowels and brief complex consonants, and variable bitrate spends bits where they are needed, giving better quality for the same file size. Constant bitrate is mainly useful when a system needs predictable data rates, such as certain streaming or broadcast chains. Either works for transcription at sensible settings.
Why do voice messages sometimes transcribe worse than recordings?
Messaging apps usually compress voice notes heavily to save data, and forwarding or re-sharing a file can compress it again. The result is often perfectly understandable to a person who knows the context but has lost much of the consonant detail a recognizer relies on. If accuracy matters, ask for the original recording rather than a forwarded message.
Should I convert a low-bitrate file to WAV before uploading?
It will not improve the transcript. Converting to WAV stores the already-degraded sound without further loss, but it cannot restore what the original compression removed, and the file becomes much larger. Upload the original file as it is. If the audio is genuinely poor, a cleaner source or a re-recording is the only real fix.