Choose a voice
Every voice speaks all 21 languages. You can play a preview in your target language before you start.
- Female
- Balanced delivery for most videos
- Female expressive
- More emotion and emphasis
- Female calm
- Soft, steady narration
- Female narrator
- Clear pace for explainers
- Female energetic
- Upbeat, lively tone
- Male
- Clear male voice
- Male calm
- Soft male narration
- Male narrator
- Steady male explainer
Timed to your video
Translations are often longer than the original — Spanish, French and German usually run longer than English. Each translated line is placed where the original was spoken and sped up by at most 1.15× when it needs to fit, with the pitch kept natural. Very dense speech can still sound brisk, so slower, clearer originals dub best.
What dubbing works best for
- Single-narrator content: explainers, tutorials, lectures, training and product demos.
- Long videos where reading subtitles would be tiring.
- Audiences who prefer dubbed content or can't read subtitles comfortably.
For interviews and conversations with several people, subtitles often work better, because one AI voice reads every speaker. Not sure? Read subtitles vs dubbing.
A 12-minute explainer dubbed into Japanese costs 600 credits (60¢), and the run also includes Japanese subtitles and both transcripts.
How it works
- Drop in a video. Your browser reads it and pulls out only the audio — the video file never leaves your device.
- The audio is processed in 30-second chunks: speech is transcribed and the spoken language is detected automatically.
- Each segment is translated into your chosen language.
- The translation is read aloud in the voice you chose, and each line is fitted to the original timing — sped up by at most 1.15× with natural pitch.
- AI vocal separation removes the original speech, and the translated voice is mixed over the remaining music and effects, which dip automatically while the new voice speaks.
- Your browser combines the original picture with the new audio and saves an MP4.
- Download the results. Audio and text on our servers are deleted after delivery, within 30 minutes at most.
What you get
- Translated video
- MP4 with your original picture and the translated voice over your original music and effects (MKV when the original video codec can't go in an MP4)
- Translated audio
- M4A with the same mix: translated voice over your original music and effects
- Transcript
- Plain text (.txt) of what was said
- Timestamped transcript
- Every line prefixed with its time, e.g. [12:04]
- Subtitles
- SRT and VTT, in the language you select
Pricing
- Transcript and subtitles
- 1 credit per minute (0.1¢), minimum 5 credits per file
- Translated voice (includes everything above)
- 50 credits per minute (5¢), minimum 2 minutes per file
- Credits
- $1 buys 1,000 credits; credits don't expire
- Free credits
- 100 credits when you sign in with Google
Good to know
- One of 8 stock voices reads the whole translation — no voice cloning and no lip-sync.
- Music and sound effects are kept by AI vocal separation, which isn't perfect: in dense mixes faint traces of the original voice can remain, and sung lyrics are removed along with the speech.
- No speaker labels — transcripts don't say who is speaking.
- Files only: no links from YouTube or social apps, and no live audio.
- Subtitles come as SRT/VTT files rather than burned into the picture.
When dubbing earns its cost, and when to skip it
Dubbing pays off when the audience listens instead of reads: viewers on phones while doing something else, people who read the target language slowly, long training sessions where a full hour of subtitles is tiring, and younger audiences. At 50 credits per minute it costs fifty times as much as subtitles, so it's worth asking whether the audience would really miss the voice.
Skip it, or test it on one video first, when the speaker's own voice is the point (a personality-led channel), when several people talk at once, when music carries the video, or when the picture is full of text that will stay untranslated anyway. In those cases translated subtitles from the subtitle generator usually serve better.
If you're unsure, trim a representative two-minute section, dub it for 100 credits (10¢) and show it to a few people from the target audience before committing a whole series.
Recording with a dub in mind
When you control the source, a few habits make the translated voice easier to follow. Speak in complete, moderately short sentences and leave a breath between them: each translated line is placed where the original was spoken, so pauses give a longer translation somewhere to go. Say product names clearly and the same way every time. Avoid jokes that depend on wordplay, and don't rely on what's in the picture to finish a sentence ("click here, then this one"). Keep background music low during speech: the dub keeps your music and effects, but a loud bed makes both recognition and the separation of voice from background harder. Producing videos that translate cleanly goes further, and what text expansion does to dubbed audio explains why some languages need more room than others.
Matching a voice to the material
- Narrated tutorials, explainers, lectures
- Female narrator or Male narrator: steady pacing that suits instruction
- Launch videos and upbeat promos
- Female energetic
- Onboarding, wellbeing or sensitive staff messages
- Female calm or Male calm
- Story-led or emotional pieces
- Female expressive
- General content with no strong tone
- Female or Male
The new voice doesn't have to match the original speaker's gender, but viewers who know a host will notice the change, so pick deliberately and keep the same voice across a series. Voice choice doesn't change the price. More on matching voice to audience in picking the right AI voice.
Reviewing a dub before it goes live
- Play the preview at 1.5× for a first pass, listening for dropped lines and obvious mistranslations.
- Slow to 0.75× for passages with names, numbers and product terms.
- Read the subtitles. In dub mode they follow what the voice actually says, including lines shortened to fit the timing, so the SRT is a fast written copy of the spoken script.
- Listen to the ends of long sentences, where a fast original leaves the translated line the least room.
- Check how names and acronyms are pronounced; when the voice says a name wrong explains where those errors start.
- Compare original and translated text side by side on the result screen for anything that reads oddly.
A fuller pre-release list is in the AI dub quality checklist.
What the dubbed file contains, and how to finish it
The translated MP4 holds your original picture, the translated voice and your original background: the original speech is removed by AI vocal separation, and the music, room tone and sound effects that remain are mixed under the new voice and lowered while it speaks. The audio is mono. Separation isn't perfect: in dense or loud mixes faint traces of the original voice can remain and the background can sound slightly thinner, and sung vocals are removed along with the speech. For most videos that's fine as delivered. If you have your edit project and want stereo music or your own levels, export a dialogue-only version and dub that; the translated audio (M4A) then comes back as essentially the voice alone, and you mix it with your clean music stem in an editor. Adding the stem to the M4A of a normal dub would double the music. Finishing an AI dub in your editor walks through it in common editors.
The same M4A is what you'd add as an extra language track on platforms that support several audio tracks per video; YouTube's multi-language audio explains how that works where it's available. If the download is an MKV because your video codec doesn't fit an MP4, check that your editor or platform accepts it before you build a timeline around it.
Costs at realistic lengths
Billed at the 2-minute minimum: 100 credits (10¢). If you have several clips from one recording, dubbing the full recording once and cutting the clips afterwards avoids paying the minimum on each.
1,250 credits ($1.25), including translated subtitles and the original-language transcript.
4,500 credits ($4.50). Long recordings are where reviewing the subtitle file instead of re-watching the whole dub saves the most time.
Each target language is its own run at the same rate.
Going further
The AI dubbing guide covers a first run step by step, and the online course use case shows dubbing across a whole curriculum. Language pages explain what changes in the target language, for example French, Korean or Hindi.
Dub video into…
Learn more
Articles that go deeper on the technology and workflows behind this tool.
- Translating interviews, panels and podcasts with more than one speaker
- Background Music in Translated Videos: What Stays and How to Perfect It
- When translations grow or shrink: what text expansion does to subtitles and dubbed audio
- Text to Speech, Explained: From Robot Voices to Neural TTS
- Inside a Synthetic Voice: The Neural TTS Pipeline Step by Step
- Chatterbox TTS: Resemble AI's Open-Source Voice Model, Explained
- AI or Human Dubbing? How to Decide for Each Project
- Dubbing, Voice-Over or Narration? The Differences That Matter
- AI Lip Sync Explained: Re-Rendering Mouths to Match New Audio
- Picking the Right AI Voice: Content, Audience, Pace and Testing
- Cloned Voices or Preset Voices? How to Decide for Dubbed Video
- What Separates a Convincing AI Voice from a Robotic One
- Why AI Dubbing Is Harder Than It Looks
- Using Synthetic Voices Without Deceiving Anyone
- Why Synthetic Voices Sometimes Stress the Wrong Word
- How One TTS Model Speaks Many Languages
- Quality-Checking an AI Dub Before It Goes Live
- How YouTube's Multi-Language Audio Tracks Work, and How to Make One
- One Voice or Many? Choosing a Dubbing Approach
- Keeping a Translated Voice in Time With the Original Speaker
- Translating Children's Videos Into Other Languages
- How to Build One Video File with Original and Translated Audio
- Finishing an AI Dub in Premiere, Resolve, Final Cut or a Free Editor
- Can Text to Speech Be Your Pronunciation Model? Uses, Limits and Shadowing
- Selling in the Buyer's Language: Translating Sales Videos
- Translating Gaming Videos: Commentary, Game Audio and Gamer Slang
- What to Do With Sponsor Reads When You Translate or Dub a Video
- How to Translate a Video Essay Without Losing Its Argument
- Loudness Normalization: LUFS, True Peak and Consistent Speech Levels
- When the Synthetic Voice Says a Name Wrong: Finding Where It Broke
- A Short History of Dubbing and Subtitles, From Title Cards to AI
- How to Improve Dialogue Clarity in Video So Everyone Can Follow the Speech
- The M&E Track: Why Every Project Should Export Music and Effects Without Dialogue
- Does Your Music License Cover the Translated Version of a Video?
- Hearing Two Voices at Once: Fixing Doubled Audio in a Video
Frequently asked questions
Does AI dubbing clone my voice?
No. You choose one of 8 stock voices. mydubly doesn't clone voices.
Is there lip-sync?
No. The dub works like a traditional voice-over; lips on camera won't match the new language.
Can I dub audio files too?
Yes — use the audio translator for podcasts and recordings.
Can different speakers get different voices?
No. One voice speaks the whole dubbed track. For interviews and panels, translated subtitles keep each speaker's own voice audible and usually work better.
Will the original voice be audible underneath the dub?
No. The original speech is separated out and removed; your music and sound effects stay underneath, lowered while the translated voice speaks. In dense or loud mixes faint traces of the original voice can remain.
Can I switch voices after the dub is made?
Choosing a different voice means running the file again, which is billed again. Play the voice previews in your target language before you start a long video.