From one voice per language to one model for many
Traditional TTS products offered voices per language. Each was recorded by a native speaker, so switching from English to French meant switching to a different person. Neural models trained on many languages at once changed that. Because they share parameters across languages, patterns learned from data-rich languages help weaker ones, and a single voice identity can be rendered in languages its reference speaker never recorded.
How text enters the model matters. Some systems convert text to phonemes first, which needs a separate pronunciation module per language. Others read characters or subword tokens directly and learn pronunciation from data, which scales to more languages but relies on the model inferring how each spelling sounds. The background on neural synthesis is in what is text to speech.
Accent leakage between languages
In training data, a speaker's voice and accent arrive bundled together. A model that learned a voice mostly from English recordings has learned English vowels, rhythm and intonation along with the timbre, and some of that can carry over when the voice speaks Spanish or Japanese. The same happens with voices conditioned on a reference clip: the reference's language colors the output.
Typical symptoms are an English-sounding r in languages that roll or tap it, vowels reduced where the target language keeps them full, stress on the wrong syllable and question contours borrowed from the source language. Mitigations include language-appropriate reference recordings, stronger language conditioning in training, and, for anyone choosing a voice, listening tests by native speakers in each target language.
Numbers, dates, units and abbreviations
Before synthesis, text has to be normalized: digits, symbols and abbreviations expanded into words. That step is language-specific, and it is where many multilingual errors start. German and many other European languages use a dot to group thousands and a comma for decimals, the reverse of English. Dates written with slashes mean different things in the United States and most of Europe. Hindi writing often groups large numbers in lakhs and crores. In Russian, Polish and Czech, number words change form depending on the noun that follows.
Suppose a German script reads "Das Update erscheint am 03/04 und kostet 1.250 Euro." A German-aware normalizer should say twelve hundred fifty euros, because the dot groups thousands, but an English-style reading would treat it as a decimal. The date could be the third of April or the fourth of March depending on the reader. Writing the date in words, "am 3. April", and keeping the amount in the target language's own conventions removes the guesswork.
Writing systems that hide pronunciation
Some scripts leave pronunciation underspecified, so the model has to infer it from context:
- Chinese characters can have more than one reading, and each syllable carries a tone that changes its meaning.
- Japanese kanji often have several readings, and the language uses pitch accent to distinguish some words.
- Arabic and Hebrew are usually written without short vowels, so the model must supply them from context.
- Russian does not mark word stress, and stress placement can change the meaning of a word.
- Hindi in Devanagari drops some inherent vowels in speech that the spelling does not show.
These languages are not worse for synthesis, but they reward a native listener's check more than languages whose spelling maps closely to sound.
Where multilingual TTS works well
- Brand consistency: every language version of a course or channel can carry the same voice.
- Coverage: less-resourced languages benefit from what the model learned elsewhere.
- Simplicity: one model and one review process instead of a separate vendor per language.
- Speed: adding a language does not require casting or booking a studio.
Limitations of multilingual TTS
- Quality is uneven across languages, generally tracking how much training data each had.
- Code-switching, where a speaker moves between languages mid-sentence, is poorly served by one tag per input.
- Names from a third language, such as a Polish surname in a Japanese sentence, are guessed.
- Tones and pitch accent may be slightly off in ways native listeners notice even when meaning is clear.
- Regional variants are covered selectively; a model may offer one accent per language or a small set.
- Accent leakage can be reduced but rarely vanishes completely.
How mydubly applies the chosen voice to the target language
In mydubly, the same eight preset voices speak all 21 supported languages. You pick a voice and a target language; the spoken language of your video is detected automatically. The default voice engine is Chatterbox Multilingual, which generates the translated sentences in the target language, with short fragments merged into whole sentences of up to about 220 characters first. The same voice reads the whole video.
Two languages offer regional voice accents: Spanish has Latin America and Spain, and Portuguese has Brazil and Portugal. The accent choice changes how the voice sounds, while the written translation uses one model per language, so vocabulary and spelling stay the same either way; translating video for regional audiences discusses what that means for a specific market. Chinese output is Simplified Chinese, Arabic output is Modern Standard Arabic and Norwegian output is Bokmål. Language pages such as Spanish and Japanese cover each language in more detail.
Testing a voice in a new language
- Pick a two-minute passage that contains numbers, a date, a name and a question.
- Dub it into the target language with AI dubbing; a two-minute file costs 100 credits, or 10 cents.
- Ask a native speaker to listen for accent, misread numbers and stress on the wrong syllable.
- Compare any problem with the translated transcript. If the number is wrong in the text, it is a translation issue, not a voice issue.
- Try a second voice on the same passage, because a voice that suits one language may suit another less well.
- Where numbers or dates misread, say them more explicitly in the original recording, for example "the third of April" rather than a bare date.
Where to go next
Once a voice passes in each target language, keep it for every language version so viewers hear one consistent presenter. Start with mydubly's AI dubbing, and if your videos switch between languages, read translating mixed-language videos before you upload.
Frequently asked questions
Does a multilingual TTS voice sound the same in every language?
The timbre usually stays recognizable, but rhythm, vowels and intonation adapt to each language, and some accent from the voice's training data can carry over. Native listeners notice the differences more than anyone else, so test each language separately.
What is accent leakage in text to speech?
It is when features of one language, usually the one a voice was learned from, show up in another: an English r in Spanish, or English stress patterns in Italian. It comes from voice and accent being learned together, and can be reduced but rarely removed entirely.
How do I stop a TTS voice from misreading numbers and dates?
Write them in the target language's own conventions, or spell them out in words when the form is ambiguous. In a dubbing workflow where you do not edit the script, say numbers and dates clearly in the original recording and check them in the translated transcript.
Can multilingual TTS handle two languages in one sentence?
Usually not well. Most models take one language tag per input, so a foreign word inside a sentence gets the main language's pronunciation. Videos that switch languages also complicate recognition, since the spoken language is detected from the audio.
What changes when I pick Spain instead of Latin America for Spanish?
In mydubly the choice changes the accent of the voice. The written translation comes from one model per language, so vocabulary and spelling do not change. If your audience expects region-specific terms, have a reviewer from that market check the translated transcript.