Preparing your files

Recording a Remote Interview over a Video Call, Cleanly and with a Backup

The most reliable way to record a remote interview is to capture audio in two places: the call platform's recording for convenience, and a local recording of each person's microphone (a double-ender) for quality. Both people should wear headphones to stop echo, get consent on the recording, and then combine the tracks into one aligned file for transcription.

8 min read · Updated

Cloud, local or double-ender recording

There are three ways to record a video call, and they differ in what they capture.

Cloud recording
the platform records the call on its servers; easy, but audio has passed through the internet connection
Local recording
your computer saves the call as you received it; same network effects, files stay with you
Double-ender
each person records their own microphone on their own device; best quality, needs combining afterwards

Cloud and local recordings capture the call after it has crossed the network. Dropouts, robotic glitches and the platform's echo cancellation and noise suppression are all baked in. That is often good enough for a transcript, and it is effortless: one click and you get a mixed file. Which options you have depends on the platform and your plan, so check its current documentation. Some platforms can also save a separate audio file per participant with local recording, which helps when you combine tracks later.

A double-ender records each voice before it touches the internet. The guest's laptop or phone records the guest; yours records you. A glitch on the call doesn't reach either file. The cost is coordination: the guest has to start a recording and send you the file. Services built for remote podcast recording automate this by recording locally in each participant's browser and uploading the files afterwards.

Headphones on both ends

If either person uses speakers, their microphone picks up the other person's voice from the speakers and sends it back. Call software tries to cancel this echo, and when both people talk at once, it often suppresses one of them, chopping off the start or end of words. That damage ends up in every recording made from the call.

Ask the guest to wear any headphones or earbuds, wired if possible. Use them yourself. It is the single most effective request you can make, and it also keeps each double-ender track clean, because each microphone hears only its own speaker.

Setting up each side of the call

For yourself, use a decent microphone close to your mouth, a quiet room with soft furnishings, and a wired network connection if you can. Record your own microphone locally in an audio app such as Audacity or a dedicated recorder. The speech recording settings article covers levels and formats.

For the guest, keep requests short and concrete, and send them a day ahead:

  • Wear headphones and sit in the quietest room available, ideally one with curtains, carpet or furniture rather than bare walls.
  • Close other apps that make sounds, and silence the phone.
  • If they have an external microphone, use it, but a laptop or phone microphone close to them is fine.
  • For a double-ender, start a voice memo on their phone placed near them, or a recording app on their computer, before the call begins.
  • Plug in the laptop, because recording and video calls drain the battery.

Do a two-minute sound check at the start of the call: listen for echo, hum, fan noise and whether the guest's voice breaks up.

Backups that save an interview

Remote recordings fail in boring ways: a guest's recording app stopped when the laptop slept, the disk filled up, the cloud recording was never started, or a cable came loose. Plan for one source to fail.

  • Run at least two recordings: the platform recording plus local double-ender tracks, or the platform recording plus a phone voice memo on your side.
  • Disable sleep on both computers for the duration, and check free disk space.
  • Ask the guest to keep their file until you have confirmed it plays all the way through.
  • After the call, write down what each file is: who, which device, start time.

If the double-ender track fails, the platform recording becomes the fallback. If the platform recording fails, the double-ender tracks are your interview.

Example: a podcast interview across time zones

Cross-border podcast, hypothetical

A host in Toronto interviews a guest in Lisbon for 52 minutes over a video call. The host records the call locally with a separate audio file per participant and her own microphone in Audacity. The guest records a voice memo on his phone, placed on a cushion next to his laptop, and wears earbuds.

Mid-interview, the guest's connection drops twice for a few seconds. The call recording has gaps and robotic patches, but his phone memo is clean throughout. After the call, he sends the M4A through a file-sharing link.

The host aligns the two local tracks in her editor using a hand clap they both made at the start, nudges the guest track by a fraction of a second near the end where the two devices' clocks had drifted, balances the levels and exports one mono WAV. The transcript costs 52 credits (5.2¢), and she keeps the call recording only as a backup.

Combining tracks into one file for transcription

  1. Collect every file and listen to the start and end of each to confirm it is complete.
  2. Import the double-ender tracks into one project in an audio editor, each on its own track.
  3. Line them up using a shared sync point: a clap at the start, or a distinctive word both tracks contain.
  4. Check alignment again near the end. Different devices record at very slightly different speeds, so long tracks can drift; shift or gently stretch the later part if they do.
  5. Balance the levels so both voices sound similar, and mute any section where one track has crosstalk or noise.
  6. Export one file, mono or stereo, as WAV or a high-quality MP3. A one-line ffmpeg alternative is ffmpeg -i host.wav -i guest.wav -filter_complex amix=inputs=2:duration=longest -ac 1 interview.wav, which lowers each input's volume, so raise the overall level afterwards if it is quiet.
  7. Play the mixed file in a few places to confirm both voices are present and nothing is doubled.

If you want a transcript you can attribute to speakers, there is another option: transcribe each person's track separately. Each transcript then contains one speaker, and you can merge them by timestamp. The catch is that long silent stretches in a single-speaker track can tempt speech recognition into inventing text, a problem explained in Whisper hallucinations, so read those transcripts carefully.

Trade-offs and things that go wrong

  • Double-enders add work. For a quick internal interview, the platform recording alone is often enough.
  • Platform noise suppression can remove more than noise, such as soft words or laughter. Some platforms let you switch to an original-sound mode, but check the effect.
  • Guests don't always follow setup instructions. Keep the list short and confirm headphones on camera.
  • Combining tracks recorded at different sample rates is fine in an editor, which converts them, but it can confuse simple scripts.
  • Two microphones in the same room, such as host and guest side by side for part of the interview, cause bleed and phasing; that case is covered in recording with multiple microphones.

Remote interviews in mydubly

mydubly transcribes and translates recorded files; it doesn't join calls or record them, and it has no live or real-time mode. Once you have a recording, upload it as audio (MP3, WAV, M4A, AAC, OGG or FLAC) or as the call's video file (MP4, MOV, WebM, MKV or M4V), up to 2 hours per file. You get a plain and timestamped transcript, SRT and VTT subtitles and, if you choose a target language, a translated transcript.

Because a file with several audio tracks is processed using only one track, export the interview as a single mixed track rather than relying on a multi-track file. mydubly doesn't label speakers, which is why the separate-tracks method above, or adding names while proofreading, is the way to attribute quotes; speaker diarization explains what automatic labeling would involve. For the editing and transcription steps after recording, see the interview transcription use case.

Next step

Write your guest instructions once and reuse them, and test the full chain, recording, combining and exporting, on a short practice call before an important interview. When the interview is in the can, upload the mixed file to audio to text.

Frequently asked questions

Is a Zoom or Teams cloud recording good enough to transcribe?

Often, yes. The audio is compressed and shaped by the call's network and echo cancellation, but speech recognition copes well with clear call audio when people wore headphones and didn't talk over each other. It falls short when the connection glitched or someone used speakers. For anything you plan to publish or quote heavily, add a local recording on each side as insurance.

What does the guest need for a double-ender?

Very little: headphones, and a second device or app that records their microphone while the call runs. A phone voice memo placed near them works well. Ask them to start recording before the call, leave it running until after you say goodbye, and send the file through a file-sharing link, since email attachments are often too small for long recordings.

How do I sync two recordings that started at different times?

Use a sync point. Ask both people to clap once, close to their microphones, right after recording starts. In your editor, line up the two clap spikes in the waveforms, and the tracks are aligned. If you forgot, use a short, distinctive word near the beginning instead. Check alignment again near the end, since device clocks can drift over long recordings.

Can I record a remote interview with just my phone?

Yes, for a voice-only interview. You can join the call on a laptop with headphones and record your own voice with a phone voice memo, while the guest does the same, which gives a simple double-ender. Recording the call audio itself on a phone depends on the operating system and app, and some don't allow it, so check what your device supports.

Should I record video or only audio?

Record video if you might publish it or need to see reactions and on-screen material. For a transcript or podcast, audio is enough and the files are much smaller. Many platforms save both anyway. If you have a video file of the call, you can transcribe it directly without extracting the audio first, as long as its format is supported.