Difficult recordings

Recording With Several Microphones: How to Mix for Transcription

Several microphones give you the cleanest possible speech, but only if you combine them carefully. Before transcribing, reduce bleed between microphones, check that no track is polarity-flipped, drop distant duplicate microphones, and export a single centered speech mix as the only audio track in the file. That one mix is what a transcription tool actually hears.

9 min read · Updated

What changes when there is more than one microphone

A single microphone gives a recognizer one view of the room. Add a second or third and every voice arrives several times: loudly in the speaker's own microphone and more quietly, slightly later, in everyone else's. How you combine those copies decides whether the transcript gets cleaner or muddier.

Done well, a multi-mic recording is about as good an input as speech recognition can get, because each person is close to a microphone. Done carelessly, the combined file sounds hollow, boomy or smeared, and the recognizer hears exactly that. This article is about the combining step: bleed, phase, separate tracks versus a mixdown, and exporting one speech track. Room echo and steady background noise are a different problem, covered in speech recognition and background noise.

Mic bleed: every voice in every microphone

Bleed, also called spill, is one person's voice picked up by another person's microphone. At a podcast table with three open microphones, each speaker is heard in all three, with the far copies quieter and carrying more room sound.

Bleed causes two kinds of trouble. The far copies add reflections to the mix, so a voice that was crisp in its own microphone gains a reverberant tail. And if you transcribe one isolated track, the other speakers still appear in it at a lower level. A recognizer may transcribe those quieter voices, transcribe them partly, or drop them, and you can't predict which.

Ways to reduce bleed at the source:

  • Use dynamic or close-talking microphones and keep each one about a hand's width from the mouth, so the wanted voice is far louder than the spill.
  • Follow the 3:1 guideline from live sound: microphones should be at least three times farther from each other than each is from its own talker.
  • Aim the dead side of a cardioid microphone toward the neighboring speaker rather than toward the wall.
  • Mute microphones that are not in use, such as a host's microphone while a recorded clip plays.

Phase and comb filtering when tracks are summed

When the same voice reaches two microphones at different distances, the copies are offset in time. Sound travels roughly 34 cm per millisecond, so a speaker 30 cm from their own microphone and 1.2 m from the next produces copies about 2.6 ms apart. Summed, the copies cancel at some frequencies and reinforce at others. This is comb filtering, and it makes speech sound thin, phasey or as if it were recorded in a tube.

Polarity is a related but separate fault. If one microphone, cable or adapter inverts the signal, or a track was flipped by mistake in the editor, summing it with another microphone that hears the same voice cancels the low end. The telltale sign is a voice that sounds full on its own track and suddenly thin when a second track is unmuted.

Mild comb filtering usually leaves speech intelligible. Strong cancellation, though, strips out the frequency detail that separates similar consonants and vowels, and it shifts as people lean in and out, so the voice keeps changing color from sentence to sentence. A good mixdown keeps one strong copy of each voice audible at any moment.

Separate tracks or a mixdown: what actually gets transcribed

Multi-mic recordings reach you in one of three shapes:

One finished mix
The recorder or mixer already combined the microphones. The balance is baked in; you can only apply EQ and cleanup to the whole thing.
A multichannel file
Several microphones stored as channels of one file, common on field recorders and on cameras that put a lavalier on one channel and the built-in microphone on the other.
Separate files or tracks
One WAV per input, or a video file carrying several audio tracks. This gives the most control.

Each shape behaves differently once uploaded. A video file can hold several audio tracks, and many tools, mydubly included, use only one of them. If the track that gets used is the camera's scratch audio rather than your mix, the transcript reflects the scratch audio. How players and tools choose a track is explained in multiple audio tracks in video files.

Channels are a separate matter. Speech recognition generally runs on mono audio, so a file with a lavalier on the left channel and a camera microphone on the right is folded into one channel before recognition. That blends the close, clean lavalier with the distant, roomy camera microphone, and any timing offset between them becomes comb filtering. Mono vs stereo for speech covers the downmix in more depth.

Transcribing each person's isolated track on its own looks like a shortcut to who-said-what. It works partly: you get one transcript per microphone and merge them by timestamp. Bleed means other voices leak into each one, though, and each extra track multiplies the review work. For most projects, one good mix plus speaker names added during review is less effort. Automatic speaker labeling is its own technology, described in what speaker diarization is.

A worked example: rebuilding a podcast mix

Example

A hypothetical three-person podcast is recorded on a four-input interface: host, two guests and a spare input left open. The interface saves each input as its own WAV and also writes a stereo mix. In the stereo mix, the spare input adds room sound, and one guest's microphone is polarity-flipped by a miswired adapter. The transcript of that mix has garbled words wherever the guest and host overlap. Rebuilding the mix from the separate WAVs, with the spare input deleted and the guest's polarity corrected, cleans up those passages without changing anything else.

The lesson is that the recorder's automatic mix is a convenience, not a master. When separate tracks exist, a ten-minute rebuild often fixes problems that no amount of cleanup on the finished mix can.

Building a single speech track, step by step

  1. Import every microphone track into a multitrack editor, such as Audacity, Reaper or the audio page of your video editor, and line them up from a common point like a clap or the first word.
  2. Solo each track and listen for faults unique to it: hum, handling noise, a dead channel or a microphone nobody used. Delete tracks that add nothing.
  3. Check polarity by playing two neighboring microphones together, inverting one, and keeping whichever setting sounds fuller. Flip a track only if the difference is clear.
  4. Where a camera or recorder captured both a close microphone and a distant one, keep the close one and remove the distant one rather than blending them.
  5. Reduce bleed between turns by lowering each microphone while its speaker is silent, using manual edits, a gate or downward expander, or an automatic mixing plugin. Make sure the first syllable of each turn survives.
  6. Balance levels so every speaker sounds roughly equally loud, then add a gentle limiter so laughter and sudden emphasis don't clip.
  7. Pan every voice to the center or export in mono. Stereo placement helps listeners but gives a recognizer nothing.
  8. Export one file: WAV or FLAC to keep, or a high-bitrate M4A or MP3 if size matters. For a video, make this mix the only audio track in the file.
  9. Before uploading, listen on headphones to one minute with crosstalk and one with laughter.

Mistakes that undo a careful multi-mic setup

  • Leaving every microphone open for the whole session just in case. Unused open microphones only add room sound.
  • Blending a camera's built-in microphone with a lavalier because it adds "air". It adds delay and echo, which smear consonants.
  • Gating so hard that the starts of words disappear. A little bleed is usually less harmful than missing syllables.
  • Exporting a stereo file with one speaker hard left and another hard right and expecting the transcript to keep them apart. The channels are combined before recognition.
  • Trusting the recorder's mix without listening. It is usually set for headphone monitoring, not for transcription.
  • Expecting a perfect mix to fix overlapping speech. When two people talk at once, both microphones contain both voices, and words will still be lost; the wider limits of crosstalk are covered in translating videos with multiple speakers.

How mydubly handles a multi-microphone file

mydubly takes one audio stream per upload. It accepts video (MP4, MOV, WebM, MKV, M4V) and audio (MP3, WAV, M4A, AAC, OGG, FLAC) up to 2 hours long. When a file has several audio tracks only one is used, so export with your speech mix as the only or main track. The browser decodes that audio to 16 kHz mono, splits it into chunks of about 30 seconds at quiet moments, and sends only compressed audio chunks over HTTPS for transcription with Whisper. The video file itself stays on your device.

You get a plain transcript, a timestamped transcript, and SRT and VTT subtitles, plus a translated transcript if you choose a target language. There are no speaker labels, so names are added by hand during review. A 50-minute three-person episode costs 50 credits (5¢) to transcribe. If you plan a dubbed version, remember that one chosen voice speaks the whole video, which is why multi-speaker shows often work better with translated subtitles.

mydubly does not mix tracks, separate voices or merge transcripts from several isolated tracks. That work belongs in your editor before upload.

Next step: test the hardest minute

Pick the minute of your recording with the most crosstalk or laughter, export it from your rebuilt mix, and run it through audio to text; a file that short costs the 5-credit minimum (0.5¢). Compare the result with the same minute from the recorder's automatic mix. If the rebuilt mix reads better, apply the same steps to the full session. For recording habits that help every file, see how to improve transcription accuracy.

Frequently asked questions

Do tracks from separate recorders need special alignment?

Usually, yes, for long sessions. Separate recorders and USB microphones each run on their own clock, so tracks that line up at the start can drift apart by a noticeable amount after an hour. Align on a clap at the start, check again near the end, and use your editor's stretch or sync feature if the end has slipped. Drift that is left uncorrected sounds like an echo between speakers.

Is one USB microphone per person a good idea?

It can work, but it is fiddly. Many computers handle several USB microphones as separate devices with separate clocks, so recording software may refuse to combine them or may introduce drift and dropouts. An audio interface or podcast recorder with several XLR inputs records all microphones on one clock, which is simpler. If you do use several USB microphones, test a full-length recording before the real session.

What sample rate should the final mix use?

Export at the rate your project already uses, typically 44.1 or 48 kHz. Speech recognition commonly runs on 16 kHz audio, and mydubly converts every file to 16 kHz mono itself, so there is no need to downsample first, and recording at a higher rate brings no recognition benefit. Audio sample rate explained covers the reasoning.

What if one guest joined over a video call?

Their voice arrives through the call's compression and sounds different from the in-room microphones. If you can, ask the guest to record a local copy on their own device and add it to your multitrack session in place of the call audio. Otherwise, keep the call audio on its own track and balance it with the others. Remote interview recording setup describes the local-copy approach.

Why does an isolated track produce text where I hear near-silence?

Bleed from other speakers can be loud enough to be partly transcribed, and recognizers sometimes produce text during long quiet passages. If you transcribe isolated tracks, listen to any passage where the text doesn't match the person whose microphone it is. Whisper hallucinations explains why quiet stretches can produce invented text.