AI voice & dubbing

What Separates a Convincing AI Voice from a Robotic One

AI voice quality comes down to four things listeners notice: whether the voice sounds human, whether every word is understood, whether the melody and emphasis fit the meaning, and whether there are audible glitches. Researchers measure it mainly with listening tests such as the Mean Opinion Score. In real projects, though, the biggest swings come from the text the model is given, especially sentence length and punctuation, and from what happens to the audio after it is generated.

7 min read · Updated

The four dimensions listeners judge

Naturalness
Does it sound like a person speaking or a machine reading? Covers voice texture, breathing and the absence of robotic monotony.
Intelligibility
Can a listener understand every word on first hearing, including names and numbers?
Prosody
Do pitch, stress, rhythm and pauses match what the sentence means?
Artifacts
Are there glitches such as buzzing, clicks, slurred syllables, or repeated and skipped words?

These dimensions overlap but are not the same thing. A voice can be perfectly intelligible and still sound unnatural, like an old announcement system. An expressive voice can sound human sentence by sentence yet keep stressing the wrong word. For dubbing, add a fifth dimension: consistency. The voice should sound like the same speaker, at the same loudness, in minute one and minute forty.

Artifacts you can hear and what usually causes them

  • Metallic or buzzy texture: usually introduced by the vocoder or decoder that turns the model's internal representation into a waveform, and more obvious on headphones.
  • Skipped, repeated or invented words: alignment failures inside the model, which become more likely with very long or very short inputs.
  • Mispronounced names and acronyms: the model guessing pronunciation from spelling, especially for words borrowed from another language.
  • Flat questions and misplaced stress: prosody prediction going wrong, often because punctuation was missing or ambiguous.
  • Drifting identity: the voice subtly changing character across a long passage or between separately generated clips.
  • Smeared consonants and a phasey sound: time-stretching pushed too far when speech is sped up or slowed down.
  • Volume jumps: clips generated separately and joined without loudness normalization.

How AI voice quality is measured

The standard measurement is the Mean Opinion Score, or MOS. A panel of listeners hears samples and rates each on a five-point scale from 1 (bad) to 5 (excellent), and the ratings are averaged per system. The method comes from telephone-quality testing, where ITU-T Recommendation P.800 describes absolute category rating, and speech synthesis research adapted it. A careful MOS study reports confidence intervals, includes real human recordings as an anchor, and uses enough listeners and sentences for differences to mean something.

Other methods fill the gaps MOS leaves:

  • Comparative tests (CMOS or A/B preference): listeners hear two versions of the same sentence and say which is better and by how much. They detect small improvements that absolute scores miss.
  • MUSHRA-style tests: listeners rate several systems side by side against a hidden reference.
  • Recognition round trips: synthesize a script, transcribe the audio with a speech recognizer and compute the word error rate against the script. Cheap and repeatable, but it rewards clarity rather than naturalness.
  • Speaker similarity: compare speaker embeddings of synthetic and reference audio, mainly used for cloned or conditioned voices.
  • Learned MOS predictors: neural networks trained to estimate listener ratings, useful for screening thousands of clips but not a replacement for human ears.

Why a published score tells you less than you think

MOS is relative to the test that produced it: which sentences, which listeners, which playback setup and which other systems were in the same session. A 4.2 in one paper and a 4.0 in another cannot be compared directly. Most tests also use short, isolated sentences read out of context, which is close to the opposite of a 30-minute dub where monotony and drift only appear after several minutes. Many evaluations are run in English, and a multilingual model can perform very differently in Korean or Finnish.

The practical answer is to treat published numbers as a filter and then test on your own content, in your own target language, with listeners who speak it.

What affects quality in a real project

  • The text: complete sentences, sensible punctuation and numbers written the way they should be read.
  • The language: models are strongest in languages that were well represented in their training data.
  • The voice: calm narrator styles tend to hide small errors; expressive styles can amplify them.
  • The reference recording behind a conditioned voice: noise or room echo in the reference can carry into every line.
  • Post-processing: tempo changes, loudness normalization and how clips are joined.
  • The upstream pipeline: in dubbing, recognition and translation errors arrive as fluent, confident speech. A perfect voice saying the wrong word is still a quality failure, which is why the AI dubbing quality checklist starts with meaning, not sound.

Why sentence length changes how a voice sounds

Neural text-to-speech models plan intonation over the whole input. Hand one a fragment and it cannot tell whether it is hearing the start of a thought or the end of one, so it may close with a falling, final-sounding pitch, or rush through as if something urgent follows. In dubbing this happens constantly, because speech recognition splits text by timing rather than grammar, and one sentence often arrives as two or three segments.

mydubly merges short fragments into whole sentences, up to about 220 characters, before synthesis, because Chatterbox Multilingual sounds better on complete sentences. The cap exists because the opposite extreme also hurts: very long inputs raise the risk of skipped words and drifting delivery.

Before and after merging

Synthesized separately, the segments "so the next thing we do" / "is open the settings" / "panel and turn on backups" each end on a falling pitch, so the line sounds like three clipped statements. Merged into "So the next thing we do is open the settings panel and turn on backups," the model hears one sentence and produces one contour with a natural pause after "do."

How mydubly protects quality after synthesis

Once each clip is generated, it is normalized: converted to mono at 24 kHz, trimmed of silence at the edges and brought to a consistent loudness, so the voice does not jump in volume between chunks. Placement is elastic. A speed-up of up to about 1.08× is treated as inaudible; beyond that, tempo can rise to 1.15× by default, with a little more allowed only when a chunk cannot otherwise fit. Lines that are too short are slowed slightly, never below 0.9×. The chunks are then joined gaplessly into one AAC track that matches the video's length. The timing logic is covered in detail in syncing translated audio with video.

Limits of today's AI voices

  • Long-form monotony: a voice that sounds lively for one sentence can feel flat after twenty minutes, because each sentence is planned largely on its own.
  • Non-verbal sounds: laughter, sighs, whispers and shouting are hard to produce on cue from plain text.
  • Rare names and mixed-language words, which the model has to guess.
  • Dense originals: when the speaker talks fast and the translation runs longer, the tempo cap means the voice sounds brisk even when nothing is technically wrong.
  • No intent: an AI voice reads what it is given. It cannot ask what you meant, so ambiguous lines get a plausible reading rather than the right one.

Run your own listening test

  1. Pick three representative passages of a minute or two: one calm, one dense with names or numbers, one emotional.
  2. Dub each with two or three candidate voices. With the 2-minute minimum per file, three short test files cost 300 credits, or 30 cents.
  3. Listen on the devices your audience uses: phone speaker, earbuds and a laptop.
  4. Ask a native speaker of the target language to score each clip from 1 to 5 for naturalness and to note every word they misheard.
  5. Compare their notes with the translated transcript to separate translation errors from voice errors.
  6. Choose the voice with the fewest distracting moments, not the one with the single most impressive sentence.

Where to go next

When a voice passes your listening test, use it for the full video with mydubly's AI dubbing and keep it for the rest of the series. If a passage keeps sounding wrong, look at the sentence itself before blaming the voice; prosody in text to speech explains how wording and punctuation steer delivery.

Frequently asked questions

What is a good MOS score for text-to-speech?

There is no universal threshold, because scores depend on the sentences, listeners and comparison systems in each test. What matters is how a system scores relative to real human recordings in the same session. Only compare numbers that come from the same study.

Why does an AI voice sound great in a demo but worse in my video?

Demos are usually short, complete sentences in the model's strongest language, picked by the vendor. A dubbed video adds fragments, names, numbers, tempo limits and twenty minutes in which small habits become noticeable. Test with your own material before deciding.

Does speeding up AI speech hurt quality?

Small changes are very hard to hear. mydubly treats a speed-up of up to about 1.08× as inaudible and caps tempo at 1.15× by default; beyond that, consonants start to smear and delivery sounds rushed. Slower, less dense originals leave more headroom.

Can poor source audio make the AI voice sound worse?

Not the voice texture, because the new voice is generated from text rather than from the original sound. Noisy audio does cause recognition errors, though, and those become fluent wrong words in the dub. Speech recognition and background noise covers what to fix before uploading.

Are female or male AI voices higher quality?

There is no general rule. It depends on the model, the training data and the specific voice. Test at least one of each on your own content, because a voice that suits a calm lecture may struggle with an energetic promo.