AI dubbing

AI Dubbing for Video

Give your video a voice in another language. Pick one of 8 AI voices, preview it in your target language, and get a dubbed MP4 with the speech fitted to your original timing.

01 — Upload

Your video

Full translation outputTimestamped transcript only — no video or audio in results.

The video file stays in this browser. Language is required: pick the spoken language for a plain transcript, or another language to translate it.

02 — Result

Your transcript

Your timestamped transcript will appear here.

Waiting for upload

25 languages available

Choose a voice

Every voice speaks all 21 languages. You can play a preview in your target language before you start.

Female
Balanced delivery for most videos
Female expressive
More emotion and emphasis
Female calm
Soft, steady narration
Female narrator
Clear pace for explainers
Female energetic
Upbeat, lively tone
Male
Clear male voice
Male calm
Soft male narration
Male narrator
Steady male explainer

Timed to your video

Translations are often longer than the original — Spanish, French and German usually run longer than English. Each translated line is placed where the original was spoken and sped up by at most 1.15× when it needs to fit, with the pitch kept natural. Very dense speech can still sound brisk, so slower, clearer originals dub best.

What dubbing works best for

  • Single-narrator content: explainers, tutorials, lectures, training and product demos.
  • Long videos where reading subtitles would be tiring.
  • Audiences who prefer dubbed content or can't read subtitles comfortably.

For interviews and conversations with several people, subtitles often work better, because one AI voice reads every speaker. Not sure? Read subtitles vs dubbing.

Example

A 12-minute explainer dubbed into Japanese costs 600 credits (60¢), and the run also includes Japanese subtitles and both transcripts.

How it works

  1. Drop in a video. Your browser reads it and pulls out only the audio — the video file never leaves your device.
  2. The audio is processed in 30-second chunks: speech is transcribed and the spoken language is detected automatically.
  3. Each segment is translated into your chosen language.
  4. The translation is read aloud in the voice you chose, and each line is fitted to the original timing — sped up by at most 1.15× with natural pitch.
  5. AI vocal separation removes the original speech, and the translated voice is mixed over the remaining music and effects, which dip automatically while the new voice speaks.
  6. Your browser combines the original picture with the new audio and saves an MP4.
  7. Download the results. Audio and text on our servers are deleted after delivery, within 30 minutes at most.

What you get

Translated video
MP4 with your original picture and the translated voice over your original music and effects (MKV when the original video codec can't go in an MP4)
Translated audio
M4A with the same mix: translated voice over your original music and effects
Transcript
Plain text (.txt) of what was said
Timestamped transcript
Every line prefixed with its time, e.g. [12:04]
Subtitles
SRT and VTT, in the language you select

Pricing

Transcript and subtitles
1 credit per minute (0.1¢), minimum 5 credits per file
Translated voice (includes everything above)
50 credits per minute (5¢), minimum 2 minutes per file
Credits
$1 buys 1,000 credits; credits don't expire
Free credits
100 credits when you sign in with Google

Good to know

  • One of 8 stock voices reads the whole translation — no voice cloning and no lip-sync.
  • Music and sound effects are kept by AI vocal separation, which isn't perfect: in dense mixes faint traces of the original voice can remain, and sung lyrics are removed along with the speech.
  • No speaker labels — transcripts don't say who is speaking.
  • Files only: no links from YouTube or social apps, and no live audio.
  • Subtitles come as SRT/VTT files rather than burned into the picture.

When dubbing earns its cost, and when to skip it

Dubbing pays off when the audience listens instead of reads: viewers on phones while doing something else, people who read the target language slowly, long training sessions where a full hour of subtitles is tiring, and younger audiences. At 50 credits per minute it costs fifty times as much as subtitles, so it's worth asking whether the audience would really miss the voice.

Skip it, or test it on one video first, when the speaker's own voice is the point (a personality-led channel), when several people talk at once, when music carries the video, or when the picture is full of text that will stay untranslated anyway. In those cases translated subtitles from the subtitle generator usually serve better.

If you're unsure, trim a representative two-minute section, dub it for 100 credits (10¢) and show it to a few people from the target audience before committing a whole series.

Recording with a dub in mind

When you control the source, a few habits make the translated voice easier to follow. Speak in complete, moderately short sentences and leave a breath between them: each translated line is placed where the original was spoken, so pauses give a longer translation somewhere to go. Say product names clearly and the same way every time. Avoid jokes that depend on wordplay, and don't rely on what's in the picture to finish a sentence ("click here, then this one"). Keep background music low during speech: the dub keeps your music and effects, but a loud bed makes both recognition and the separation of voice from background harder. Producing videos that translate cleanly goes further, and what text expansion does to dubbed audio explains why some languages need more room than others.

Matching a voice to the material

Narrated tutorials, explainers, lectures
Female narrator or Male narrator: steady pacing that suits instruction
Launch videos and upbeat promos
Female energetic
Onboarding, wellbeing or sensitive staff messages
Female calm or Male calm
Story-led or emotional pieces
Female expressive
General content with no strong tone
Female or Male

The new voice doesn't have to match the original speaker's gender, but viewers who know a host will notice the change, so pick deliberately and keep the same voice across a series. Voice choice doesn't change the price. More on matching voice to audience in picking the right AI voice.

Reviewing a dub before it goes live

  1. Play the preview at 1.5× for a first pass, listening for dropped lines and obvious mistranslations.
  2. Slow to 0.75× for passages with names, numbers and product terms.
  3. Read the subtitles. In dub mode they follow what the voice actually says, including lines shortened to fit the timing, so the SRT is a fast written copy of the spoken script.
  4. Listen to the ends of long sentences, where a fast original leaves the translated line the least room.
  5. Check how names and acronyms are pronounced; when the voice says a name wrong explains where those errors start.
  6. Compare original and translated text side by side on the result screen for anything that reads oddly.

A fuller pre-release list is in the AI dub quality checklist.

What the dubbed file contains, and how to finish it

The translated MP4 holds your original picture, the translated voice and your original background: the original speech is removed by AI vocal separation, and the music, room tone and sound effects that remain are mixed under the new voice and lowered while it speaks. The audio is mono. Separation isn't perfect: in dense or loud mixes faint traces of the original voice can remain and the background can sound slightly thinner, and sung vocals are removed along with the speech. For most videos that's fine as delivered. If you have your edit project and want stereo music or your own levels, export a dialogue-only version and dub that; the translated audio (M4A) then comes back as essentially the voice alone, and you mix it with your clean music stem in an editor. Adding the stem to the M4A of a normal dub would double the music. Finishing an AI dub in your editor walks through it in common editors.

The same M4A is what you'd add as an extra language track on platforms that support several audio tracks per video; YouTube's multi-language audio explains how that works where it's available. If the download is an MKV because your video codec doesn't fit an MP4, check that your editor or platform accepts it before you build a timeline around it.

Costs at realistic lengths

45-second social clip

Billed at the 2-minute minimum: 100 credits (10¢). If you have several clips from one recording, dubbing the full recording once and cutting the clips afterwards avoids paying the minimum on each.

25-minute course module

1,250 credits ($1.25), including translated subtitles and the original-language transcript.

90-minute conference recording

4,500 credits ($4.50). Long recordings are where reviewing the subtitle file instead of re-watching the whole dub saves the most time.

Each target language is its own run at the same rate.

Going further

The AI dubbing guide covers a first run step by step, and the online course use case shows dubbing across a whole curriculum. Language pages explain what changes in the target language, for example French, Korean or Hindi.

Dub video into…

Learn more

Articles that go deeper on the technology and workflows behind this tool.

Frequently asked questions

Does AI dubbing clone my voice?

No. You choose one of 8 stock voices. mydubly doesn't clone voices.

Is there lip-sync?

No. The dub works like a traditional voice-over; lips on camera won't match the new language.

Can I dub audio files too?

Yes — use the audio translator for podcasts and recordings.

Can different speakers get different voices?

No. One voice speaks the whole dubbed track. For interviews and panels, translated subtitles keep each speaker's own voice audible and usually work better.

Will the original voice be audible underneath the dub?

No. The original speech is separated out and removed; your music and sound effects stay underneath, lowered while the translated voice speaks. In dense or loud mixes faint traces of the original voice can remain.

Can I switch voices after the dub is made?

Choosing a different voice means running the file again, which is billed again. Play the voice previews in your target language before you start a long video.