AI video translation

Code-Switching, Bilingual Speakers and Videos That Change Language

If you translate video with mixed languages using a tool that detects the spoken language automatically, expect inconsistency where the language changes: a stretch of audio is usually treated as being in one language, so the minority language can be transcribed phonetically, translated at the wrong stage or skipped. When the switches are long and clean, such as an English keynote followed by Spanish questions, splitting the file at the change and processing each part separately fixes most problems. When speakers switch mid-sentence, subtitles with human review are the more reliable route.

6 min read · Updated

What counts as a mixed-language video

Mixed-language content comes in several shapes, and they behave very differently in a translation pipeline:

  • Block switches: a conference session where the keynote is in one language and the Q&A in another, or a documentary that alternates between narrated sections and interviews in a second language.
  • Turn-by-turn switches: an interview where the host asks in English and the guest answers in Hindi, with each turn lasting a minute or so.
  • Code-switching: one speaker moving between languages within a sentence, common among bilingual speakers, for example Spanish and English in the same breath.
  • Loanwords and brand names: an otherwise monolingual video sprinkled with foreign terms, like English software names in a German tutorial.

The last case is rarely a problem. The first is easy to fix. The middle two are where automatic tools struggle.

How automatic language detection handles mixed audio

Multilingual speech recognition models such as Whisper start by predicting which language they are hearing, then transcribe with that language in mind. The prediction is made for a window of audio, not word by word. Long recordings are processed in pieces, so the language is effectively identified from each stretch of audio rather than once for the whole file.

That works well when each stretch is in one language. When a window contains two languages, the model has to pick one, and the other language is handled as if it were the chosen one. Closely related languages, short utterances and heavy accents make the guess less certain. The mechanics of identification are explained in spoken language identification.

What goes wrong in the transcript and the translation

When a stretch of audio is identified as language A but contains language B, a few things can happen, and you may see different ones in the same video:

  • Language B is written out phonetically in language A's spelling or script, producing words that look plausible but mean nothing.
  • Language B is rendered directly into language A during recognition, so the transcript reads smoothly but no longer reflects what was said.
  • Short phrases in language B are dropped entirely.
  • Translation then treats everything as language A. Segments already in the target language may be paraphrased or passed through unchanged, so the dub can switch register or repeat itself awkwardly.

The danger is not obvious errors but quiet ones. A transcript that reads fluently can still hide a minute of the guest's answer that was summarized or lost. That is why mixed-language videos need a listen-through, not just a read-through.

When mixed-language content works fine

Not every mixed video needs special handling. Automatic processing is usually reliable when:

  • One language clearly dominates and the other appears only as names, terms or brief quotes.
  • Switches happen at long, clean boundaries, such as a change of speaker after several minutes.
  • The secondary language is the same as your target language and you only need subtitles, so viewers can hear the original for those parts anyway.

If your video fits one of these, process it normally and spend review time on the moments where the language changes.

Splitting the file at language changes

For block switches, splitting turns one hard problem into two easy ones. Each part is monolingual, so detection is reliable and the transcript reflects what was said.

  1. Scrub through the video and note the timestamps where the language changes for longer than a minute or so.
  2. Cut the file at those points. A lossless cutter, or ffmpeg with stream copy, splits without re-encoding, though cuts land on the nearest keyframe; an editor export works too.
  3. Decide per part what you need. A part already in the target language may need nothing, or only same-language captions.
  4. Process each remaining part in the video translator, choosing the same target language and voice for every part.
  5. Reassemble the translated parts in an editor in the original order, and merge or shift the subtitle files to match the full timeline.

Check the per-file minimums before splitting into many small pieces: a translated voice track is billed at a minimum of 2 minutes per file, and transcripts at a minimum of 5 credits per file.

Worked example

Suppose a 30-minute conference recording has an 18-minute keynote in Spanish followed by 12 minutes of Q&A in English, and you want an English version. Processed as one file, the Q&A segments risk being re-rendered or paraphrased. Split at the 18-minute mark, only the keynote needs a dub: 18 × 50 = 900 credits, or 90¢. The English Q&A stays as recorded, and a transcript-only pass for English captions on that part costs 12 credits. For a German version, both parts are dubbed separately for 900 + 600 = 1,500 credits, or $1.50, the same price as one file because both parts exceed the minimum.

Other workarounds for harder cases

  • For turn-by-turn interviews, use translated subtitles rather than a voice track, and have a bilingual reviewer correct each answer against the audio using the timestamps.
  • Run transcript mode first to see where detection went wrong before paying for a dub.
  • Re-record a short narrated summary of a code-switched passage in one language when the exact wording matters less than the content.
  • For research interviews, budget for human transcription of the code-switched sections; research interview transcription discusses review workflows.

Limits that splitting can't solve

Splitting only helps when switches are long enough to cut around. Intra-sentence code-switching cannot be separated with a razor tool, and cutting every 20 seconds multiplies per-file minimums and editing time. Accents that blur the line between related languages, such as regional varieties of Hindi and Urdu, can still confuse identification inside a single part. And if part of the audio is in a language outside the ones a tool supports, recognition for that part is unreliable however you cut it. In those cases, plan for human review or human transcription of the affected sections.

Mixed-language video in mydubly

mydubly detects the spoken language automatically; you only choose the target language. Recognition runs on Whisper over audio windows of about 30 seconds, so a recording that changes language can be transcribed and translated inconsistently around the switch. mydubly supports 21 languages, and both parts of a split file must be in one of them for reliable results.

Practical consequences: split block switches into separate files, use transcript mode to inspect a mixed recording cheaply before dubbing, and remember that one chosen voice reads the whole output, so a dubbed bilingual interview becomes one voice in one language. Transcripts and SRT or VTT subtitles in both the spoken and target language come with every full run, which gives a bilingual reviewer what they need to check each switch.

Where to go next

Start by classifying your video: loanwords, block switches, turn-by-turn or true code-switching. For the first two, process it in the video translator, splitting at block boundaries. For the last two, generate subtitles and plan a bilingual review. If the speakers are also overlapping, read translating video with multiple speakers before deciding between voice and subtitles.

Frequently asked questions

Can I tell mydubly which language is spoken in my video?

No. The spoken language is detected automatically from the audio and you only pick the target language. For a recording that changes language, splitting it into single-language files is the way to make detection reliable.

Why did part of my bilingual video come out in the wrong language in the transcript?

The recognition model decided that stretch of audio was in the other language and transcribed it accordingly, either phonetically or by rendering it into that language. Check that section against the audio and, if the switch is long, process it as its own file.

Does code-switching affect subtitles as well as dubbing?

Yes, because both start from the same transcript. Subtitles are easier to fix, though: a reviewer can edit the cue text in the SRT file, whereas a dubbed voice would need regenerating.

Is it cheaper to split a mixed-language video into parts?

It costs the same when every part is longer than the per-file minimum, and less when some parts do not need translating at all. Many tiny parts can cost more, because each file has a 2-minute minimum for a voice track.

How should I handle English product names in a non-English video?

Usually no special handling is needed, because a few loanwords do not change the detected language. Check how the names were spelled in the transcript, since recognition may adapt them to the surrounding language's spelling.