Why the format question matters less than people expect
A transcription engine does not listen to your file the way an audio engineer does. Most speech recognition systems, including Whisper, convert incoming audio to a single channel at 16 kHz before they analyze it. That is a much narrower slice of sound than a 44.1 kHz stereo WAV or a 192 kb/s MP3 carries. Much of the detail that separates a lossless file from a good lossy one sits above the range the model ever sees.
What the model does depend on is the voice itself: how close the microphone was, how much the room echoed, whether music or traffic sat underneath, and whether the level was clipped. A clean voice memo saved as M4A will usually transcribe better than a WAV recorded across a noisy room. If you want to improve results, the recording settings for speech matter more than the save format.
What actually differs between the four formats
- WAV
- Usually uncompressed PCM audio. Every sample is stored as recorded, so files are large and nothing is lost.
- FLAC
- Lossless compression. The file is smaller than WAV, yet it decodes back to exactly the same samples.
- MP3
- Lossy compression. A perceptual model discards sound it predicts you won't notice; the bitrate controls how much is thrown away.
- M4A
- A container, most often holding AAC (lossy) and sometimes Apple Lossless. Phone voice recorders commonly save AAC in M4A.
The useful split is lossless versus lossy, not WAV versus MP3. WAV and FLAC are interchangeable in quality, and so are well-encoded MP3 and AAC at similar settings. If you want the mechanics behind this, the explainer on what an audio codec is covers how lossy encoders decide what to discard.
When compression starts to hurt recognition
Lossy compression becomes a problem for speech in a few specific situations rather than as a general rule.
- Very low bitrates. Encoders at the bottom of their range tend to blur consonants such as s, f and th, which are exactly the sounds that separate similar words. The article on audio bitrate for speech recognition goes into where that floor sits.
- Generation loss. Each time a lossy file is decoded and re-encoded, new artifacts are added on top of the old ones. A recording that went from a phone to a messaging app to an editor and back out as MP3 may have been compressed three or four times.
- Compression stacked on a weak signal. A faint voice under background noise leaves the encoder little to work with, and compression artifacts can make the remaining speech harder to separate from the noise.
Converting a lossy file to WAV or FLAC does not undo any of this. The new file is bigger, but it contains the same damaged audio, so there is no transcription benefit in converting an MP3 to WAV before uploading it.
Size, storage and handling
Uncompressed audio is large. A minute of CD-quality stereo WAV (44.1 kHz, 16-bit) takes a little over 10 MB, so a two-hour recording runs past a gigabyte. The same two hours as a 128 kb/s MP3 comes to roughly 115 MB. FLAC lands between them; how much it saves depends on the material.
Size matters when you email files, sync them to cloud storage or send them to a service that uploads the whole thing. It matters less for transcription itself, since the engine downsamples anyway. For long-term storage of interviews, oral histories or podcast masters, size is usually a price worth paying for a lossless copy you can re-edit later.
Choosing a format by situation
There is no single right answer, so match the format to the job.
- You only need the text, once: use whatever the recorder produced. An M4A from a phone or an MP3 from a podcast host is fine as is.
- You will edit the audio before or after transcribing: keep a WAV or FLAC master, edit that, and export a lossy copy only for distribution.
- You are archiving research or oral history recordings: keep the lossless original. FLAC saves space without giving up anything, and WAV is the most universally readable. Check your archive's own deposit requirements.
- You only have an old, heavily compressed file: transcribe it as it is. Converting it won't help; cleaning the audio in an editor sometimes does.
- You are short on storage or bandwidth: a moderate-bitrate MP3 or AAC file is a reasonable working copy for transcription, provided you keep a better original somewhere.
A researcher records an interview on a field recorder as 48 kHz WAV, about 500 MB. She transcribes the WAV directly, files it in the project archive as FLAC, and sends a 96 kb/s MP3 to her co-author for listening. All three copies would give essentially the same transcript; only the WAV and FLAC are suitable for later editing or deposit.
A sensible workflow for recordings you will transcribe
- Record in the best format your device offers. On a field recorder that is usually WAV; on a phone the default AAC in M4A is fine for speech.
- Copy the original off the device before doing anything else, and leave that copy untouched.
- If you need to trim, denoise or level the audio, do it on a lossless copy and export once.
- Transcribe from the original or the single edited export, not from a file that has been forwarded through chat apps.
- Keep the lossless master and the transcript together, named so the link between them is obvious.
- Make lossy copies only for sharing, and treat them as disposable.
Mistakes and trade-offs to watch
- Converting MP3 to WAV "for better accuracy" costs disk space and changes nothing.
- Saving the only copy of an important interview as a low-bitrate MP3 to save space removes the option to edit or restore it later.
- Stereo files with the speaker on one channel and silence on the other are fine for archiving but worth checking before transcription; see mono vs stereo for speech for how channels are combined.
- WAV files with unusual sample formats can fail in some players and apps. If a WAV won't open, re-export it as standard 16-bit or 24-bit PCM.
- Lossless does not mean clean. A WAV of a voice in an echoing hall is still an echoing hall.
How mydubly treats MP3, WAV, FLAC and M4A
mydubly's audio to text tool accepts MP3, WAV, M4A, AAC, OGG and FLAC, as well as video files, up to two hours per file. Whatever you upload, the browser decodes it on your device to 16 kHz mono, splits it into chunks of about 30 seconds at quiet moments, and compresses each chunk to Opus at about 32 kb/s before sending it over HTTPS for transcription with Whisper. The original file itself stays on your device.
In practice this means a large WAV does not cost you a large upload, and a lossless file has no transcription advantage over a decent MP3 of the same recording. Decoding a very long uncompressed file does use more of your device's memory, so on a phone a compressed copy can be easier to handle. The WAV to text and MP3 to text pages list format specifics.
You get a plain transcript, a timestamped transcript, and SRT and VTT subtitle files. The price is the same for every format: 1 credit per minute with a 5-credit minimum, so a 45-minute interview costs 45 credits (4.5¢). mydubly does not label speakers, so multi-person recordings need speaker names added by hand.
The upload side of this trade-off is covered in reducing upload size for video translation, and how speech to text works explains why the recognizer ends up working from 16 kHz mono either way.
Keep the master, transcribe the convenient copy
If you remember one rule, make it this: archive lossless, transcribe whatever is closest to the original recording, and never convert a lossy file to lossless expecting an improvement. When you are ready, open the audio to text tool with your file and compare the transcript against a few minutes of the recording to judge its quality.
Frequently asked questions
Is 16-bit or 24-bit WAV better for transcription?
For transcription, the difference does not matter. Bit depth affects how much quiet detail and headroom a recording has, which is valuable while recording and editing, especially if levels were set low. Once audio is reduced to 16 kHz mono for recognition, a properly leveled 16-bit file carries everything the model uses. Record in 24-bit if your recorder offers it, for the editing headroom, and transcribe either version.
What MP3 bitrate is safe for speech recognition?
Mono speech at moderate bitrates, such as 64 to 128 kb/s, generally keeps the consonant detail recognition needs. Problems tend to appear at the very bottom of the range, especially when the sample rate has also been reduced. If a file sounds noticeably swishy or underwater when you listen on headphones, expect more transcription errors on unclear words.
Is OGG or Opus a good choice for recordings I want to transcribe?
Yes, at sensible bitrates. Opus was designed with speech in mind and holds up well at low bitrates, which is why many voice apps use it. As with MP3, avoid re-encoding it repeatedly and keep a lossless original if you plan to edit. The Opus codec explainer covers its strengths and limits.
Should I convert an iPhone voice memo before transcribing it?
No. Voice memos are usually AAC in an M4A container, which transcription tools generally accept directly. Converting it to WAV or MP3 adds a step and, in the MP3 case, another round of lossy compression. Upload the original M4A, ideally copied straight from the phone rather than forwarded through a messaging app that may recompress it.
Does a lossless file give better timestamps?
Not in any meaningful way. Timestamps come from where the recognizer places speech segments, and that depends on clear pauses and clean speech rather than on lossless storage. A file with heavy compression artifacts can blur word boundaries slightly, but a decent lossy file and a lossless one should give timestamps that line up with each other closely.