What voice cloning actually does
Modern voice cloning comes in two broad forms. Zero-shot cloning conditions a pretrained text-to-speech model on a short reference clip, and the model imitates the timbre and speaking style it hears without any extra training. Fine-tuned cloning goes further: the model's weights are adapted on a larger set of recordings of one speaker, which tends to capture more of their habits, such as how they end sentences or where they breathe.
Either way, the output is speech that never existed as a recording. The system has learned how the person sounds, not what they would say or how they would say it, so the words, emphasis and pacing still come from the script and from the model's own predictions. That distinction matters in dubbing: a cloned voice reading a translation is one person's timbre attached to delivery decisions that person never made.
If you want the mechanics of how a reference recording steers a neural model, see how AI voice generation works.
What makes a voice a stock voice
A stock or preset voice is a fixed synthetic voice offered to every user of a product. Under the hood it may be produced exactly the way a clone is, by conditioning a model on a reference recording, but the provider chooses and manages that reference, and users select the voice from a list instead of supplying audio.
That arrangement has practical consequences:
- Everyone using the product hears the same voice, so it is predictable from one video to the next.
- You do not need permission from a particular person to use it.
- The provider can test the voice in each supported language before anyone publishes with it.
- The voice is not exclusive: another channel may be using the same one.
Voice cloning vs stock voices at a glance
- Who it sounds like
- Cloning: a specific person whose audio you provide. Stock: a preset voice chosen by the provider.
- What you need
- Cloning: clean reference audio plus the speaker's informed consent. Stock: nothing beyond picking from a list.
- Audience recognition
- Cloning: people who know the speaker hear continuity. Stock: neutral, nobody expects it to be someone in particular.
- Cross-language behavior
- Cloning: the original speaker's accent can leak into the target language. Stock: can be checked by the provider in every language it offers.
- Misuse exposure
- Cloning: impersonation and fraud risk if the audio or model escapes. Stock: much lower, because no real person's identity is being copied by the user.
- Series consistency
- Cloning: depends on the quality of each reference clip. Stock: the same voice every time.
When a cloned voice is worth it
Cloning earns its complexity when the voice itself is part of what the audience comes for. A few situations where that holds:
- Personality-led channels where viewers subscribe to a host, and a stranger's voice in the dubbed version would feel like a different show.
- Founder or executive messages where the speaker's identity carries authority, and the speaker has agreed in writing to have their voice synthesized.
- Long-running series with a recognizable narrator, where changing voices between languages would break the brand.
- Voice banking, where a person who expects to lose their voice records samples so a synthetic version can speak for them later.
In each case the speaker is identifiable, available and willing. If any of those three is missing, a clone becomes hard to justify.
Suppose a software founder records a 10-minute product update in English every week and wants a Spanish version. Viewers know her voice, so a clone might feel more personal. But the update is mostly feature walkthroughs, the screen carries most of the information, and she would have to approve synthetic speech in a language she does not read. A calm stock narrator at 500 credits (50 cents) per episode delivers the same information without asking her to sign off on words she cannot check.
Risks and limits of cloning
Consent is the first issue, and it is not a one-time checkbox. A person can agree to a clone reading product updates and still object to it reading an apology, an advertisement or a political message. Good practice is to write down what the clone may be used for, for how long, and how the person can withdraw. The wider questions are covered in AI voice ethics.
Misuse is the second. A convincing clone can impersonate someone in a phone scam, a fake endorsement or a fabricated statement. Whoever holds a cloned voice model, or the clean recordings behind it, is holding something worth stealing.
Quality is the third, and it surprises people. A clone reproduces timbre more reliably than delivery, so it can sound like the person while phrasing sentences in ways they never would. Across languages, the reference speaker's vowels and rhythm can carry into the new language, so a voice cloned from an English recording may speak French with an English tinge. Listeners who know the real person are also the harshest judges: a small slip in a familiar voice reads as uncanny, while the same slip in an unfamiliar stock voice passes unnoticed.
Finally, platforms and a growing number of jurisdictions expect realistic synthetic versions of real people to be disclosed, and some uses are banned outright. Check the current rules of every platform you publish on before releasing a cloned voice.
When stock voices are the better choice
Stock voices fit content where information matters more than identity: training modules, tutorials, lectures, product documentation, internal announcements and explainers. They also suit videos where several people speak. A clone of one participant reading everyone's lines would put other people's words in that person's mouth, while a neutral narrator signals that what you hear is a translation. For more on that trade-off, see single voice vs multi voice dubbing.
A stock voice is also the safer default when the original speaker is unavailable, has left the organization or has died. Nobody can be misrepresented by a voice that was never theirs.
How mydubly handles voices
mydubly's AI dubbing offers eight preset voices: Female (balanced), Female expressive, Female calm, Female narrator, Female energetic, Male, Male calm and Male narrator. There is no voice cloning, and users cannot upload a recording of their own voice or anyone else's.
The default voice engine is Chatterbox Multilingual, an open-source text-to-speech model from Resemble AI. Some of the presets are produced by conditioning that model on a reference recording, the same technique cloning products rely on, but the references are fixed by mydubly rather than supplied by users. The model itself is described in Chatterbox TTS.
One chosen voice is used for the whole video, and the translated MP4 replaces the original speech with that voice while keeping the original music and effects underneath. Spanish and Portuguese each offer two voice accents, Latin America or Spain and Brazil or Portugal, which changes how the voice sounds without changing the written translation.
Choosing between them in practice
- Ask whether your audience knows the speaker's voice and would notice its absence. If not, a stock voice is enough.
- If they would, confirm the speaker is willing, available to review output and comfortable with the specific uses you plan.
- Check whether the content has several speakers; if it does, a neutral voice usually reads more honestly than one person's clone.
- Dub a short test clip in the target language and play it to a native speaker of that language.
- Pick one voice and keep it for the whole series so returning viewers hear continuity. How to choose an AI voice covers matching a voice to tone and pace.
Where to go next
If a preset voice fits your project, open mydubly's AI dubbing tool, choose the target language and one of the eight voices, and run a two-minute test first, which costs 100 credits or 10 cents. If you decide a clone is essential, use a provider that documents its consent process, and keep a written record of exactly what the speaker approved.
Frequently asked questions
Can a stock voice be based on a real person's recording?
Often, yes. Many preset voices are created by conditioning a model on a reference recording or by training on a voice actor's sessions. The difference is that the provider selects and manages that recording, so as a user you are not copying anyone's identity.
How much audio does voice cloning need?
Zero-shot systems can imitate a voice from a short clip, while fine-tuned clones typically use far more recorded speech. Requirements vary widely between providers, so check their documentation. Clean audio without music, noise or room echo matters more than sheer length.
Will a cloned voice sound like me in another language?
It will usually carry your timbre, but the pronunciation, rhythm and intonation in the new language come from the model. Many cloned voices keep traces of the original language's accent, which some listeners find charming and others find distracting.
Can I use my own voice for a mydubly dub?
No. mydubly offers eight preset voices and does not accept voice uploads. If you need your own voice, you can run a translated transcript with the video to text tool and record that script yourself or hand it to a voice actor.
Do I need to disclose that a dub uses a stock AI voice?
Disclosure rules focus mostly on synthetic content that could be mistaken for a real person, but a short note such as 'Spanish audio generated with an AI voice' is good practice either way. Check each platform's current policy, and see the acceptable use policy for what mydubly permits.