The parts of prosody
- Intonation
- The pitch contour across a phrase: falling at the end of a statement, often rising on a yes-or-no question in English, rising and falling through a list.
- Stress and prominence
- Which syllable in a word, and which word in a sentence, stands out through pitch, loudness and length.
- Rhythm and tempo
- The timing pattern of syllables and the overall speaking rate. English squeezes unstressed syllables; Spanish and French give syllables more even timing.
- Pauses and phrasing
- Where speech breaks into chunks, and for how long.
- Voice quality
- Breathiness, tension and energy, a large part of how emotion is heard.
Prosody carries meaning the words alone do not. "You're coming." and "You're coming?" differ only in pitch. "I didn't say she took the money" changes meaning depending on which word is stressed: stress "I" and someone else said it; stress "took" and perhaps she borrowed it. A listener decodes all of that without thinking, and a synthetic voice has to encode it without being told.
How a TTS model decides on prosody
Earlier neural systems predicted prosody explicitly. Models in the FastSpeech family, for example, estimate a duration, pitch and energy value for each sound before generating audio. Newer systems that generate speech as sequences of learned audio tokens handle prosody implicitly: intonation emerges from patterns learned across thousands of hours of recordings.
Either way, the model's evidence is limited to the text it receives and, for conditioned voices, a reference recording that sets the general style. It has no access to the speaker's intent, the previous scene, the picture or the joke that set up this line. So it chooses the most probable reading, which is usually a neutral one: broad focus, gentle prominence near the end of the sentence, a fall at the full stop. For a deeper look at the architecture, see how AI voice generation works.
Why translated scripts lose emphasis
In dubbing, emphasis disappears in two places. First, transcription records words, not stress. A speech recognizer returns text with punctuation and casing; the fact that the speaker leaned hard on one word is gone. Second, languages mark focus differently. English relies heavily on pitch accent, so the same word order can carry contrast through voice alone. Spanish and Italian often move the focused element toward the end of the sentence, French tends to use cleft constructions such as "c'est elle qui", and Japanese uses particles to mark topic and subject. A translation engine working from unmarked text usually produces neutral word order, so the contrast is lost before the voice ever sees the line.
A presenter says "We didn't cut the price, we cut the setup time," hitting "price" and "setup time" hard. The neutral Spanish "No redujimos el precio, redujimos el tiempo de configuración" is accurate, but a synthetic voice may stress the verbs and blur the contrast. Rewritten as "Lo que redujimos no fue el precio, sino el tiempo de configuración," the structure itself carries the contrast, so even a neutral reading lands it.
Context limits make this worse. Translation and synthesis both work on limited windows of text, so emphasis that depends on something said a minute earlier is invisible to them. Context in machine translation explains the translation side of that problem.
How punctuation steers synthetic delivery
Punctuation is the main control a plain-text TTS model responds to. Behavior varies between models, but these tendencies are common:
- Full stop: a falling, final contour and a clear pause.
- Comma: a short break, often with a slight continuation rise. Missing commas produce breathless run-ons.
- Question mark: a rise on yes-or-no questions in many languages. Questions starting with who, what or why often fall even in human speech, and models differ in how they treat them.
- Exclamation mark: more energy and higher pitch. Used on every sentence, it sounds shrill.
- Colon and dash: an anticipatory pause, handled inconsistently from model to model.
- Ellipsis: hesitation or trailing off; some models insert a long silence, others ignore it.
- Capital letters for emphasis: unreliable, and some models read an all-caps word as an acronym and spell it out.
Because these are tendencies rather than rules, test the exact punctuation habits you plan to use on the voice you plan to use.
Writing lines that read well aloud
- Keep one idea per sentence, and aim for sentences a person could say in one breath.
- Put the new or important information at the end of the sentence, where neutral intonation naturally places prominence.
- Make contrast explicit in the words, with structures such as "not X but Y", "only" or "even", instead of relying on vocal stress.
- Punctuate for breathing: add commas where a speaker would pause, even if grammar does not demand them.
- Write numbers, dates and abbreviations the way they should be spoken whenever the written form is ambiguous.
- Watch for words whose pronunciation depends on meaning, such as read, lead, live and record, and rephrase when context is thin.
- Read every line aloud yourself. If you stumble, the model probably will too.
Where better prosody makes the biggest difference
Getting prosody right pays off most where intonation carries structure. Tutorials with steps and lists depend on the listener hearing where one item ends and the next begins. Quizzes and course material contain questions that must sound like questions. Marketing lines often hinge on a contrast, as in the example above. Long lectures benefit because varied rhythm and clear phrasing reduce listener fatigue over forty minutes. In each of these, small wording changes in the source can be worth more than switching voices.
What you cannot control through text
- Exact pitch and timing targets. Some cloud TTS services accept SSML, a W3C markup standard with tags for emphasis, breaks and prosody, but many neural models, especially open-source ones, ignore it or support only part of it.
- Emotional arcs across many sentences, because each input is planned largely on its own.
- Run-to-run variation: models that sample during generation can read the same sentence slightly differently each time.
- What happens after synthesis: tempo adjustment to fit a time slot compresses pauses along with words.
- Truly ambiguous lines, where only knowledge of the situation reveals the intended stress.
Prosody in mydubly dubs
mydubly generates its voice track from the translated transcript automatically, so there is no script to mark up by hand. Several parts of the pipeline exist to protect prosody anyway. Short recognition fragments are merged into whole sentences, up to about 220 characters, before synthesis, so Chatterbox Multilingual can plan one contour over a full sentence instead of guessing at pieces. Timing fit keeps tempo changes small: up to about 1.08× is treated as inaudible, the default ceiling is 1.15×, and short lines are slowed slightly but not below 0.9×.
The voice you choose sets the baseline delivery. Female expressive carries more emphasis than Female calm, and the narrator voices keep a steadier explanatory pace. The full list is on the AI dubbing page.
The biggest lever, though, is the original recording. The punctuation in the transcript comes from how the speaker sounded: clear pauses and complete sentences produce full stops and commas, which survive translation and steer the voice. Explicit contrast in the spoken wording survives too. Making videos translation-ready covers recording habits that help.
Where to go next
Dub a short passage that contains a question, a list and a contrast, and listen for whether each one survives; with mydubly's AI dubbing a two-minute test costs 100 credits. If delivery is fine but something else sounds off, AI voice quality covers the other dimensions listeners judge.
Frequently asked questions
What is the difference between prosody and intonation?
Intonation is one component of prosody: the rise and fall of pitch across a phrase. Prosody also includes stress, rhythm, tempo, pauses and voice quality. A voice can have reasonable intonation and still sound off because its pauses or stress are wrong.
Can I use SSML to control prosody in a dub?
Some cloud text-to-speech services accept SSML tags for emphasis, pauses and pitch, but many neural models ignore them. mydubly generates the voice from the translated transcript automatically, so there is no SSML input; the levers are voice choice and how the original is spoken.
Why does an AI voice sound monotone on long videos?
Most models plan delivery sentence by sentence, so there is no overarching arc across minutes of speech. Similar sentence lengths make it worse. Varied sentence structure in the source and a more expressive voice style both help.
Why do questions sometimes come out sounding like statements?
Often the question mark never reached the model, because the speaker's rising intonation was transcribed as a full stop. Questions that start with who, what or why also tend to fall naturally. Check the transcript's punctuation on lines that matter.
Does speeding up a line change its prosody?
A tempo change compresses words and pauses together while keeping the overall contour, so small adjustments sound natural. Larger ones squeeze pauses until phrasing blurs and the line sounds hurried, which is why dubbing tools cap how far they speed speech up.