What happens to several speakers in an AI translation
Speech recognition hears a single mixed soundtrack. It transcribes whatever words it can make out, in order, with timestamps, but unless a separate speaker diarization step is added, nothing in the output records which person spoke which segment. Translation then works on that unlabeled text, and the voice model reads it. The result keeps the conversation's content and timing while flattening its cast. What is speaker diarization explains the technique that some tools use to separate speakers, and why it has its own error modes.
One voice for everyone: what viewers actually hear
In a single-voice dub, the same synthetic voice asks the question and answers it. Each line still lands where the original speaker's line was, so the rhythm of the conversation survives, and viewers can usually follow turn-taking from the picture: who is on camera, whose mouth is moving, who is nodding. Problems start when the picture doesn't help. Audio-led video podcasts, voices off camera, quick back-and-forth and two people with similar framing all leave the viewer guessing.
Because a dub removes the original speech, the real voices disappear entirely, including the cues that normally identify speakers, such as pitch, accent and laughter. The single voice vs multi-voice dubbing article weighs that trade-off in general terms.
Transcripts and subtitles without speaker labels
Unlabeled transcripts are fine for searching and quoting but awkward for reading a conversation. Subtitles are less affected, because they appear while the original audio plays and viewers can hear who is speaking. If you want labels, add them yourself: subtitle files are plain text, so a reviewer can prefix lines with names or initials, or use the common convention of starting a line with a dash when the speaker changes within a two-line cue.
Crosstalk, interruptions and overlapping speech
Overlap is the main accuracy problem in multi-speaker video. When two people talk at once, a recognizer typically follows the louder voice, drops the quieter one, or merges fragments of both into one sentence that neither person said. Short backchannel responses such as 'right' and 'mm-hm' are often dropped. Each of those problems carries into translation and then into the voice track, where a merged sentence is read as a single smooth line.
Original audio: the guest says 'The budget was cut in March, so we' as the host cuts in with 'Wait, March this year?' The transcript may come out as 'The budget was cut in March, so we wait, March this year?' Translated and voiced, it becomes one confusing sentence in one voice. With subtitles over the original audio, a viewer at least hears the interruption and can work out what happened.
For new recordings, separate microphones and a moderator who discourages talking over each other make a big difference. For existing recordings, there is nothing to separate after the fact, so plan for review around the busiest exchanges.
When translated subtitles serve multi-speaker video better
Subtitles keep the original voices audible, which solves the identity problem without any extra work. They are the stronger choice for:
- Panels and roundtables with three or more speakers, or frequent interruptions.
- Audio-led video podcasts where speakers are not always visible.
- Documentary interviews, where a person's own voice and emotion are part of the story.
- Debates and Q&A sessions, where it matters who said what.
- Any content where a misattributed statement could cause harm.
Subtitles and transcripts are also the cheapest output, at 1 credit per minute in mydubly; the subtitle generator produces them without a voice track.
When one translated voice works, and how to help it
A single voice works well for host-and-guest interviews with clear turns, explainer videos where one person talks most of the time, and training videos with a presenter and occasional questions. A few habits help:
- Choose a neutral narrator-style voice rather than an energetic one, so the voice reads as an interpreter rather than as one of the people on screen.
- Tell viewers at the start, in the description or a title card, that the translation uses one voice for all speakers.
- Publish the translated subtitles alongside the dub, with speaker names added where turns are hard to follow.
- Keep the original-language version available, so viewers who want the real voices can find them.
How to choose an AI voice covers voice selection in more depth.
A workaround for two speakers: two runs, two voices, one edit
For a two-person interview where distinct voices really matter, you can build a two-voice dub yourself in a video editor.
- Translate the video twice into the same language, once with a female voice and once with a male voice (or two contrasting styles).
- Download the translated audio file from each run.
- In your editor, place the original video with both translated audio files on separate tracks underneath, aligned at the start.
- Mute or cut each track wherever the other speaker is talking, switching at the pauses between turns.
- Listen at every switch. Both runs are timed against the same original speech, so lines should land close to the same moments, but the two voices' lines differ slightly in length, and the wording can vary a little between runs.
- Export the result, and fix any wording differences in the subtitle file so it matches what viewers hear.
Suppose a journalist wants a Spanish version of a 30-minute interview between a female host and a male guest. Two dubbed runs cost 1,500 credits each, 3,000 credits or $3.00 in total. Cutting between the tracks at about 60 turn changes takes an editor perhaps an hour, and gives viewers a different voice for each person. For a panel of five, this approach stops being practical, and subtitles are the better route.
Where multi-speaker translation falls short
Even with care, AI translation of conversations has limits worth stating plainly. Overlapping speech loses words. Speaker identity is lost in the voice track unless you rebuild it by hand. Grammatical gender is a quiet trap: in languages such as French, Spanish, Hebrew or Polish, a speaker's own sentences often mark their gender, and a translator working without speaker information may choose inconsistently, especially when a man and a woman are talking. Laughter, sighs and tone of voice, which carry a lot of meaning in conversation, are not reproduced by the voice track. And humor between speakers, the teasing and callbacks, is among the hardest content to translate well.
mydubly and multi-speaker video
In mydubly, one chosen voice from the 8 stock voices reads the entire video, and transcripts do not label speakers. There is no voice cloning, so the translated voice will not sound like any of the original speakers, and the translated MP4 removes the original speech rather than mixing the new voice over it, keeping only the original music and effects underneath. What you do get from each run is everything the workarounds above need: the translated MP4, the translated audio file on its own, SRT and VTT subtitles you can label, and transcripts in both languages for review. For conversation-heavy formats, the podcast translation and conference talk translation pages describe typical setups.
Choosing your approach
Count the speakers and watch for overlap. Two speakers with clean turns: a single-voice dub is usually fine, and the two-run edit is there if you need distinct voices. Three or more, or lots of crosstalk: lead with translated subtitles. When you are ready to hear how a single voice handles your conversation, try a short excerpt with AI dubbing before committing the full recording.
Frequently asked questions
Can AI give each speaker in my video a different translated voice?
Some tools attempt it with speaker separation, but mydubly uses one chosen voice for the whole video. For two speakers you can approximate it by translating twice with different voices and cutting between the two audio tracks in an editor.
Why does my translated interview transcript not show who is speaking?
Speech recognition outputs words and timestamps, not identities. Labeling speakers needs a separate diarization step, which mydubly does not perform, so add names to the subtitle or transcript files during review if you need them.
What happens to the parts where people talk over each other?
The recognizer usually captures the louder voice and may drop or merge the other, so overlapping passages are where most errors appear. Review those moments first, and consider subtitles for content with frequent interruptions.
Is a single-voice dub acceptable for a video podcast?
It can be if speakers take clear turns and are usually on camera. For audio-led episodes with off-camera voices or fast banter, translated subtitles over the original audio are easier to follow.
Which voice style suits a multi-speaker translation?
A calm or narrator-style voice tends to work, because it sounds like an interpreter relaying everyone rather than like one of the participants. Test it on a short excerpt with a busy exchange before choosing.