What a sample rate measures
Sound reaches a microphone as a continuously changing air pressure. To store it digitally, a converter measures that pressure at regular intervals and records each measurement as a number. The sample rate is how often it measures. At 48 kHz, there are 48,000 samples for every second of sound, and each channel of a stereo recording gets its own 48,000.
Sample rate is one of three settings that describe uncompressed audio. Bit depth sets how precisely each sample is stored, which affects the noise floor and how much headroom you have before clipping. Channels set whether there is one signal or several. Bitrate, which you see on compressed files such as MP3 or AAC, is a different thing again: the size of the encoded stream after compression. Choosing bit depth and other recorder settings is covered in audio recording settings for speech.
Nyquist and aliasing in plain terms
The sampling theorem associated with Harry Nyquist and Claude Shannon says that a sample rate can faithfully represent frequencies up to half its value. That limit is called the Nyquist frequency: 8 kHz for 16 kHz audio, 22.05 kHz for 44.1 kHz, 24 kHz for 48 kHz.
Frequencies above the limit do not simply vanish. If they reach the converter, they fold back and appear as false tones at lower frequencies, an artefact called aliasing. Every competent converter and resampler therefore applies a low-pass filter first, removing content above the Nyquist frequency before sampling or reducing the rate. That filter is why well-made sample-rate conversion sounds clean, and why crude conversion can add faint whistles or a gritty edge.
Where the common rates come from
- 8 kHz
- Traditional telephone audio. Captures up to 4 kHz; phone networks historically passed roughly 300 to 3,400 Hz
- 16 kHz
- Wideband speech, used by many VoIP systems and by speech recognition models such as Whisper. Captures up to 8 kHz
- 22.05 and 32 kHz
- Older multimedia and broadcast uses; still seen in some recorders and archives
- 44.1 kHz
- The audio CD standard, also common for music and podcasts
- 48 kHz
- The standard for video, film and broadcast. Most cameras, phones and editors use it for video sound
- 88.2 to 192 kHz
- High rates used in music production and some archival work; they leave room for processing but add nothing audible to speech
The odd-looking 44.1 kHz has a historical reason: early digital audio was stored using video equipment, and the rate fitted neatly into video line structures of the time. Video production settled on 48 kHz, which is why mixing music sources and video sources so often means converting between the two.
How much bandwidth speech actually needs
The fundamental pitch of adult voices sits low, typically somewhere around 85 to 255 Hz, but intelligibility depends on much higher frequencies. Vowels are distinguished by resonances called formants, which extend up to around 3 to 4 kHz. Consonants such as s, f, th and sh carry energy well above that, often beyond 4 kHz and up past 8 kHz.
That explains the familiar weakness of phone calls. At 8 kHz sampling, nothing above 4 kHz survives, so s and f sound alike and people spell names letter by letter. At 16 kHz, the band up to 8 kHz includes most of the consonant detail, which is enough for both listeners and recognizers. Higher rates add air and sparkle that matter for music and for pleasant listening, but very little that changes which words are understood. Narrowband recordings get their own article, transcribing phone call recordings.
Why recognition models standardize on 16 kHz
Speech recognition models are trained on audio at one fixed rate, and every input is converted to that rate before the model sees it. Whisper uses 16 kHz. The stages that follow, turning audio into a spectrogram and reading it with a neural network, are explained in how speech to text works.
The practical consequence is reassuring: feeding a recognizer 48 kHz audio does not confuse it, because it is resampled down first, and the information above 8 kHz is simply not used. Recording at a higher rate is therefore never a mistake for transcription. What does matter is that the content below 8 kHz is clean, which depends on the microphone, the room and the distance far more than on any sample-rate setting.
Converting between sample rates
Converting audio from one rate to another, called resampling, is routine. Going down, the resampler low-pass filters and then keeps fewer samples. Going up, it calculates new samples between the existing ones. Two points are worth remembering:
- Upsampling adds no information. An 8 kHz phone recording converted to 48 kHz still has nothing above 4 kHz. A spectrogram view in an audio editor shows this instantly as an empty region above the original limit.
- Mismatches are worse than conversions. If audio recorded at 48 kHz is played or imported as if it were 44.1 kHz without conversion, it plays slower and lower in pitch; the opposite mismatch makes it faster and higher. In video editing, an uncorrected mismatch can also show up as audio drifting away from the picture.
A field recorder captures a 10-minute interview at 48 kHz, 24-bit, stereo: 48,000 samples × 3 bytes × 2 channels is 288,000 bytes per second, about 173 MB in total. Converted to 16 kHz, 16-bit mono, the same interview is 32,000 bytes per second, about 19 MB. For a transcript, both versions produce essentially the same input, because the recognizer would resample to 16 kHz anyway. For the podcast edit, the 48 kHz original is the one to keep.
Choosing and checking the sample rate
- For any recording that might be edited or published, record at 48 kHz, or at 44.1 kHz if your audio-only workflow already uses it.
- Match the project rate in your editor to the rate of most of your material, normally 48 kHz for video.
- Check what a file actually contains with a media inspector, or in an audio editor's file information panel, before exporting.
- Leave resampling for recognition to the transcription tool rather than exporting a 16 kHz copy yourself; it saves a step and keeps your original intact.
- If audio sounds too fast, too slow or high-pitched, suspect a rate mismatch first and fix the setting rather than the audio.
Sample rate mistakes, and what a higher rate can't fix
- Recording at 96 or 192 kHz to improve a transcript. It produces larger files without changing what the recognizer hears.
- Upsampling phone or old voicemail audio in the hope of better clarity. The missing frequencies stay missing.
- Mixing 44.1 and 48 kHz files in a project without letting the editor convert them, which can cause pitch, speed or drift problems.
- Blaming the sample rate for a muffled or noisy recording. Those faults come from the microphone, placement and room; see the guide to improving transcription accuracy.
- Exporting at a low rate such as 11 or 22 kHz to save space. Lossy compression at a sensible bitrate saves far more space while keeping the full speech band.
What mydubly does with your sample rate
mydubly accepts recordings at whatever rate they were made, within its supported formats. In your browser, the audio is decoded once and converted to 16 kHz mono, the format its Whisper-based recognition expects, before being split into chunks of about 30 seconds and compressed for upload; the article on reducing upload size covers the size side of that. A 48 kHz camera file and a 44.1 kHz podcast export therefore both reach the recognizer in the same form, and you do not need to convert anything beforehand.
What the conversion cannot do is add bandwidth that was never captured. An 8 kHz call recording is accepted and transcribed, but its missing high frequencies make some consonants harder to tell apart, so names and numbers deserve an extra check. The original file stays on your device either way.
Next step
Look up the sample rate of your usual recordings and set your recorder and editor to 48 kHz if they are not already. Then transcribe a typical file with audio to text, or a video with video to text. A 10-minute test costs 10 credits (1¢), which is enough to judge whether the recording, not the rate, needs attention.
Frequently asked questions
Should I use 44.1 kHz or 48 kHz?
Use 48 kHz for anything involving video, because cameras, editors and video platforms are built around it and you avoid conversions. Use 44.1 kHz if you work only with audio and your existing music or podcast workflow is already set to it. For speech the audible difference between the two is negligible; consistency within a project matters more than which one you pick.
Should I convert my recording to 16 kHz before transcribing it?
There is no need. Transcription tools that use 16 kHz models resample the audio themselves, normally with good filtering. Converting yourself adds a step, risks a poor-quality resample if the tool you use is crude, and leaves you with a lower-rate copy you might later mistake for the original. Keep the recording at its native rate.
Is 8 kHz audio good enough for transcription?
It usually works, but with less margin. At 8 kHz nothing above 4 kHz is captured, so consonants such as s and f become harder to distinguish, and names, numbers and unfamiliar terms suffer most. Clear, close speech still transcribes reasonably well. Review those parts against the audio, and where you can, record future calls with a wideband or higher-rate method.
Why does my audio sound like chipmunks or slowed down?
That is the classic sign of a sample-rate mismatch: the samples were recorded at one rate and played or imported as if they were another, so everything is sped up or slowed down along with the pitch. Check the file's real rate in a media inspector and set the project or import rate to match, or let the editor convert the file properly.
Does a higher sample rate mean better quality?
Only up to the point where the extra bandwidth matters. Above 44.1 or 48 kHz, the gain is in headroom for heavy processing in music production, not in anything listeners hear in speech. For transcription, everything above 8 kHz is discarded during resampling. Microphone quality, distance and room noise have a far larger effect on how good a voice recording sounds.