How your voice is copied
There is no separate voice recording to make and no voice to choose. mydubly works out who speaks when, then builds a sample for each person from their own speech in your file: up to about 10 seconds of their clearest lines, with music and background noise removed. That sample is what the voice model copies when it speaks the translated lines. Someone with less than about 3 seconds of clean speech can't be copied reliably, so they get the copied voice of a similar-sounding person in the same file, or otherwise a stock voice matching their voice (female or male).
The samples are working files. They are deleted with the job's audio and transcripts after delivery, within 30 minutes at most, and nothing is used to train models.
What carries over and what changes
- Carries over
- The character of the voice: pitch range, timbre and the overall impression of who is speaking
- Follows the new language
- Pronunciation, rhythm and intonation of the target language
- Follows your video
- Timing: each line is placed where the original was spoken and sped up by at most 1.15× when the translation runs long
- Varies
- How much of a speaker's accent survives, and how closely shouting, whispering or strong emotion are reproduced
Calm, clear speech copies best. A presenter talking to camera in a quiet room is the ideal case; a voice heard mostly over loud music or across a noisy room gives the model less to work with. The article on voice cloning vs stock voices explains the trade-offs between the two approaches.
Several people, several voices
Interviews, podcasts on video, panels and family videos usually have more than one speaker. Up to 8 people per file are told apart and each is dubbed in a copy of their own voice, so a host and a guest stay distinct in the new language. Telling people apart is automatic: quick crosstalk, one-word interjections or two very similar voices can put a line in the wrong voice, and overlapping speech is voiced one line after another. Translating videos with multiple speakers covers how to record and check them.
Getting a better copy of your voice
- Start with a few seconds of you speaking alone, without music underneath; an intro to camera is ideal.
- Record with a microphone close to your mouth. A lavalier or a USB microphone beats a laptop microphone across the desk.
- Keep background music low under speech, or add it after translation in your editor.
- Avoid rooms with strong echo; reverb gets copied into the voice along with you.
- Dub a two-minute clip first. At the two-minute billing minimum it costs 100 credits (10¢), and you'll hear how your voice sounds in the new language before translating a long video.
The guide to recording audio for speech has settings for common microphones and apps.
Consent and honest use
Dub only voices you have the right to use: your own, or people who have agreed to be dubbed. Voice copies must not be used to impersonate anyone, put words in someone's mouth or mislead viewers about who said what; the acceptable use policy sets out the rules. When you publish a dubbed version, saying that the translation and voice are AI-generated is good practice, and some platforms ask for it. The ethics of AI voices discusses consent in more depth.
What it doesn't do
- No lip-sync: the picture is unchanged, so lips on camera don't match the new language.
- No voice library: every voice comes from the people in your file.
- No translated voice in Czech or Vietnamese; those two get subtitles and translated transcripts.
- Files only: no links, no live audio, and up to 2 hours per file.
What a dub costs
Dubbing into Spanish costs 500 credits (50¢) and returns the Spanish MP4 in the presenter's copied voice, a Spanish M4A, both transcripts and Spanish subtitles.
Dubbing into German costs 2,100 credits ($2.10). Host and guest each speak German in a copy of their own voice.
Dubbing costs 50 credits per minute, with a 2-minute minimum per file. There is no subscription: you pay per minute in credits, which don't expire. For a language-specific walkthrough, see English to Hindi video translation or Spanish to English.
Using the tool on this page
- The spoken language is detected automatically; you select the output language — the spoken language itself for an untranslated transcript, or another language for a translation. Files can be up to 2 hours long.
- Only the audio is sent for processing — the video stays on your device — and audio and text are deleted within 30 minutes of delivery (how files are protected).
- You can download a translated video that keeps your music and effects, the translated audio as M4A, a transcript and subtitles (SRT and VTT). Every output, step and limit is explained on the AI dubbing page.
Frequently asked questions
Do I need to record a separate voice sample?
No. The sample for each person is taken from their own speech in the video you translate. A few seconds of clear, solo speech somewhere in the file is enough to make a copy.
Which languages can my voice speak?
19: Arabic, Chinese, Dutch, English, Finnish, French, German, Hebrew, Hindi, Italian, Japanese, Korean, Norwegian, Polish, Portuguese, Russian, Spanish, Swedish and Turkish. Czech and Vietnamese are available as subtitles and transcripts.
Is my voice stored or used for training?
No. Voice samples are deleted with the rest of the job's files after delivery, within 30 minutes at most, and nothing is used to train models.
Can I dub someone else's voice?
Only with their permission, and never to impersonate or mislead. Interviews and videos with guests are fine when the people in them have agreed.