What mono and stereo really mean
A mono recording has one channel: one signal that plays identically from every speaker. A stereo recording has two channels, left and right, which can carry different signals so sounds appear positioned between the speakers.
In practice, files labelled stereo come in three quite different forms, and telling them apart matters more than the label:
- True stereo
- Two related signals from a stereo microphone or mix, with sounds spread across the image. Common in music, film ambience and field recordings
- Dual mono
- Two independent mono signals in the left and right channels, such as host on one side and guest on the other. Common on interview recorders and call recorders
- Duplicated mono
- The same single microphone copied into both channels. Sounds identical to mono but takes twice the space
- One-sided
- Speech in one channel and silence or noise in the other, usually from a mono microphone plugged into one input of a stereo recorder
Separate audio tracks in a video file are a different thing from channels within one track; that topic has its own article on videos with multiple audio tracks.
When stereo helps speech
Stereo is useful for speech in a few specific situations:
- Two speakers on separate channels. Recording each microphone to its own side keeps them apart, so you can rebalance levels later or process each voice separately.
- Safety recordings. Some recorders write a second channel at lower gain, giving you a clean copy if the main channel clips.
- Film and documentary sound, where room tone and ambience give a sense of place around the dialogue.
- Speech mixed with music, which usually benefits from a stereo image for listeners.
For one person speaking into one microphone, stereo adds nothing a listener can hear. A mono export also means a lossy codec spends its whole bitrate on that one signal rather than splitting it, as explained in audio bitrate for speech recognition.
How stereo becomes mono for recognition
Speech recognition models take one channel of audio, so any multichannel file is downmixed first. The usual method is simple: the channels are added together and scaled, which for stereo means roughly averaging left and right sample by sample. Surround mixes are folded down the same way, with the center channel, where film dialogue usually sits, included in the result.
Averaging works well when both channels carry the same speech, and reasonably well for dual mono, where both voices end up audible in one signal. Two situations go wrong.
The first is one-sided audio. If the voice is only on the left and the right is silent, averaging halves the voice level, roughly a 6 dB drop. The words survive and recognition is usually fine, just quieter. If the empty channel carries hiss or hum instead of silence, though, that noise is mixed in at the same weight as the voice.
The second is phase cancellation, and it can be dramatic.
Phase problems and other stereo pitfalls
Sound waves have peaks and troughs. If the right channel carries the same signal as the left but upside down, a fault called inverted polarity, adding the two makes peaks meet troughs and the voice cancels almost completely. In stereo it sounds wide and strange; in mono it may nearly vanish, leaving only thin, hollow remnants. Miswired cables, adapter faults and some recorder settings cause this.
Milder versions are more common:
- Two microphones at different distances from the same speaker capture the voice a few milliseconds apart. Summed to mono, some frequencies cancel and others reinforce, giving a hollow, phasey sound called comb filtering.
- Stereo-widening effects and some podcast plugins create width by manipulating phase, which can collapse badly in mono.
- Mixed crosstalk. On dual mono interview recordings, each microphone also picks up the other speaker faintly and slightly delayed, which adds some comb filtering when both channels are summed.
The cure for multi-microphone setups, from placement to gain to mixing, is covered in mixing multiple microphones for transcription. Here the important thing is to notice the problem before it reaches a transcript.
A filmmaker records an interview with an external stereo microphone. Through headphones the voice sounds oddly spacious, as if it comes from inside the head. A transcript of the file comes back patchy, with whole phrases missing. Switching playback to mono reveals the cause: the voice almost disappears. One channel's polarity was inverted by a faulty adapter. Inverting that channel back in an audio editor and exporting a mono file restores a full, clear voice, and the next transcript is complete.
Checking a stereo file before transcription
- Listen in mono. Most operating systems have a mono audio setting under accessibility options, and many audio editors and players have a mono button. If the voice gets thinner, quieter or hollow, investigate.
- Look at the waveform in an audio editor. Two channels that look like mirror images of each other indicate inverted polarity; one flat channel indicates a one-sided recording.
- If a phase correlation meter is available, readings near +1 mean the channels agree; readings near -1 mean they cancel.
- Fix the file. Invert one channel to correct polarity, keep only the good channel for one-sided audio, or export a properly mixed mono file from your editor.
- For a quick command-line fix, ffmpeg can keep just the left channel with -af pan=mono|c0=c0, or average both with -ac 1, while copying any video stream unchanged.
How mydubly handles stereo and mono files
mydubly decodes the soundtrack in your browser to a single 16 kHz mono channel before splitting it into chunks of about 30 seconds and sending compressed speech for recognition with Whisper. Stereo, dual mono and surround files are all combined into that one channel, so you do not need to convert them, but a file that cancels in mono will transcribe poorly, and a one-sided recording will reach the recognizer at a lower level. Checking in mono first, as above, tells you which case you have.
Because everything is mixed into one channel, mydubly does not use channel separation to tell speakers apart, and its transcripts do not label speakers. If your recording has each person on their own channel, you can take advantage of that yourself: export each channel as its own mono file and transcribe them separately. Each transcript then contains one voice with its own timestamps, which you can merge by time. That costs one job per file, with a minimum of 5 credits each. Speaker labelling as a technique is explained in what speaker diarization is.
For dubbing, one chosen voice speaks the whole translated video, and the translated audio is mono, including the original music and ambience that are kept under the voice. If you want music or ambience in stereo around it, export a dialogue-only version from your project and dub that, then mix the returned translated voice with your clean stereo music stem in an editor. Adding the stem to the audio of a normal dub would double the music.
Where mono is the safer default
- Solo narration, voice memos, lectures and single-speaker tutorials.
- Phone and lavalier recordings, which capture one signal regardless of the file format.
- Files destined only for transcription, where the recognizer uses one channel anyway.
- Archival copies of speech, where mono halves the size of uncompressed files without losing anything.
Choose stereo, or keep it, when you have two independent microphones you may want to rebalance, ambience that matters to the final film, or music in the mix. For multi-person video, translating videos with multiple speakers covers the trade-offs beyond channels.
Next step
Play your usual recording format in mono once and listen for anything that thins out or disappears. If it holds up, upload it as it is to audio to text; if each speaker is on their own channel, consider the split-channel approach above. The interview transcription page shows how a two-person recording typically moves from recorder to transcript.
Frequently asked questions
Should I record my voice in mono or stereo?
Record one voice with one microphone in mono. A single microphone produces a single signal, so a stereo file would just hold two copies of it, doubling the size without any audible benefit. Use stereo when you have two microphones you want kept apart, or when ambience and space are part of the production, as in film or field recording.
Does stereo improve transcription accuracy?
Not by itself. Recognition models work on one channel, so stereo is mixed down before transcription. A stereo file can even transcribe worse than mono if its channels cancel each other or if one channel is mostly noise. What improves accuracy is a clean, close voice; the channel layout matters only in that it must survive the mixdown intact.
Why does my audio sound hollow or disappear in mono?
That points to a phase problem. If one channel has inverted polarity, summing the two cancels most of the voice. If two microphones caught the same voice at slightly different times, summing creates comb filtering, which sounds hollow. Invert one channel or keep only the better microphone, then export mono and listen again before transcribing.
My recording only has sound in the left channel. Is that a problem?
Usually not for transcription. When the channels are averaged, the voice simply ends up quieter, which recognizers generally handle. It is worth fixing if the empty channel carries hiss or hum, because that noise is mixed in, or if you are publishing the audio, because listeners will hear the voice in one ear. Keep the good channel and export mono.
What is the difference between dual mono and stereo?
Dual mono holds two independent mono signals, such as two separate microphones, in the left and right channels. True stereo holds two related signals that together create a sense of position and space. Both are stored as two-channel files, so the label alone does not tell you which you have. Listening on headphones, or looking at the two waveforms, does.