What text to speech actually does
A TTS system takes a string of characters and produces an audio waveform of someone saying it. That sounds simple, but writing leaves out a great deal of what speech carries. Text has no explicit stress, no pitch contour, no pauses beyond punctuation, and many words that are spelled one way and spoken another. The system has to infer all of that before it can make a sound.
Every TTS engine, old or new, therefore solves two problems. First it works out what to say: expanding "Dr." into "doctor" or "drive", reading "1999" as a year rather than a quantity, and deciding where the sentence rises and falls. Then it works out how to sound: producing audio with the right timbre, rhythm and loudness for a particular voice. How those two problems are split up, and how much is learned from data rather than written by hand, is what separates the generations of the technology.
Three generations of synthetic speech
Machines that speak are older than computers. Bell Labs demonstrated the Voder, a keyboard-operated speech synthesizer, at the 1939 New York World's Fair. Software TTS took off from the 1970s onward, and the field has moved through roughly three approaches since.
- Rule-based (formant) synthesis
- Hand-written rules drive an electronic model of the vocal tract. Very intelligible and tiny, but unmistakably robotic. Classic screen readers and the DECtalk family of the 1980s worked this way.
- Concatenative (unit selection)
- A voice actor records many hours of speech, which is cut into small units. At runtime the engine stitches together the units that best match the target sentence. Natural on sentences close to the recordings, audibly glitchy at the joins elsewhere.
- Statistical parametric
- Models such as HMM-based systems predict acoustic parameters that a vocoder turns into sound. Smooth and flexible, small enough for devices, but often muffled or buzzy.
- Neural
- Deep networks learn the mapping from text to sound directly from recorded speech. WaveNet (2016) showed neural audio generation could sound strikingly human, and Tacotron-style models (2017) learned pronunciation and prosody end to end. Most voices you hear today are neural.
Within the neural era there has been a further shift. Earlier neural systems predicted a spectrogram and handed it to a separate vocoder. Newer ones often treat speech like language: they convert audio into discrete tokens and train a model to predict those tokens from text, which makes it easier to imitate a voice from a short sample. The architecture details are covered in how AI voice generation works.
Where you already hear text to speech
TTS is so common that it is easy to forget it is there. Screen readers such as VoiceOver, TalkBack and NVDA read interfaces aloud for blind and low-vision users, and many of them still favor fast, crisp voices over natural ones because users listen at very high speeds. Navigation apps, phone menus, smart speakers and public announcements rely on it. E-readers and browsers can read articles aloud, and language-learning apps use it to model pronunciation.
Video is a newer frontier. Creators use synthetic narration for explainers and short-form clips, and AI dubbing chains speech recognition, translation and TTS to give an existing video a voice in another language. That last use asks more of the technology than reading a menu does, because the synthetic speech has to fit timing that a human set.
What modern neural TTS does well
Today's best neural voices have genuine strengths that make them usable for long-form listening, not just short prompts.
- Natural timbre: the breathiness, resonance and micro-variation of a real voice rather than a buzzing tone.
- Stable pronunciation of common vocabulary across long passages, without the fatigue or drift a human reader shows after an hour.
- Plausible sentence-level prosody: questions rise, lists have rhythm, and commas produce believable pauses.
- Many languages from one model, which matters for translation workflows.
- Speed and repeatability: the same text can be re-voiced in minutes after an edit, which is why synthetic narration suits content that changes often.
Where neural TTS still falls short
The weaknesses are just as real, and knowing them helps you decide when synthetic speech is good enough.
- Ambiguous text: heteronyms like "read", "lead" or "live" are resolved from context, and the model sometimes guesses wrong. Abbreviations, symbols and unusual number formats trip it up too.
- Names and jargon: brand names, people's names and technical terms outside the training data get plausible but wrong pronunciations.
- Acting: TTS reads; it does not perform. Sarcasm, comic timing, grief or a deliberate hesitation rarely come through unless the model is given explicit control and someone directs it.
- Emphasis: a human knows which word in a sentence carries the meaning. A model infers it from wording and punctuation, so emphasis can land on the wrong word, especially in translated text.
- Occasional glitches: neural models can skip or repeat a word, mumble an ending, or produce an odd artifact on very short or very long inputs.
- Consistency over very long output: tone can wander slightly between sentences generated separately.
Take the line "I didn't say she stole the money." A human narrator chooses which word to stress, and each choice changes the meaning. A TTS engine given only the text will usually produce a neutral reading with mild stress near the end. If your script depends on that kind of emphasis, rewrite it so the meaning is carried by the words, for example "I never accused her; someone else did."
How to judge a text to speech voice before relying on it
A short structured test tells you more than a demo page does.
- Pick a passage of about a minute that represents your real content, including at least one name, one number and one question.
- Generate it and listen once without reading along, the way your audience will.
- Listen again while reading the script and note every mispronunciation, wrong stress and odd pause.
- Listen to a long stretch, ten minutes or more if you can, to check for fatigue in the listener rather than the voice: monotony shows up over time.
- Decide which problems you can fix by rewriting the text (spelling out numbers, splitting long sentences) and which mean you need a different voice or a human reader.
Text to speech inside mydubly
mydubly uses TTS as the final voice step of video and audio translation. Your speech is transcribed with Whisper, translated by a neural machine translation engine, and then voiced in the target language by the default voice engine, Chatterbox Multilingual, an open-source neural model from Resemble AI. You can read more about that model in our explainer on Chatterbox TTS.
Because TTS sounds best on complete sentences, mydubly merges short transcript fragments into whole sentences of up to about 220 characters before synthesis, then fits each line back into the original timing. You pick one of 8 stock voices for the whole video; there is no voice cloning, and the picture is never altered. Voicing costs 50 credits per minute with a 2-minute minimum, so a 10-minute explainer is 500 credits, or $0.50, and the same run also returns SRT and VTT subtitles plus transcripts in both languages. The AI dubbing page shows the voices and output files, and the audio translator does the same for podcasts and recordings.
Where to go next
If you want to hear what neural TTS sounds like on your own material, the most useful test is a short clip from a real video. Open AI dubbing, pick a target language and a voice, and listen critically using the checklist above. For help picking among voices, see how to choose an AI voice.
Frequently asked questions
Is text to speech the same as an AI voice?
Most AI voices today are produced by neural text-to-speech models, so in everyday use the terms overlap. Strictly, TTS is the task of turning text into speech, and an AI voice is one particular voice a model produces. Voice conversion systems, which turn one person's speech into another voice, are a related but different technology.
Why do screen reader voices still sound robotic?
Many screen reader users listen at very high speeds, and older formant-style voices stay intelligible when sped up far beyond natural pace. They are also small and fast to respond. Natural-sounding neural voices are available on most platforms, but plenty of experienced users prefer the crisp, predictable older voices.
Can text to speech pronounce names correctly?
Common names usually come out fine, but uncommon personal names, brand names and borrowed words are frequent errors because the model guesses from spelling. Some engines accept phonetic hints; otherwise respelling a name phonetically in the script is the usual workaround.
How long can a text to speech recording be?
There is no practical limit on total length, because long scripts are generated sentence by sentence or in chunks and joined. The quality concerns are consistency between chunks and listener fatigue on a flat delivery rather than any hard cap. In translation tools the limit is usually the file length the tool accepts, which is 2 hours per file in mydubly.
Is synthetic speech good enough for a professional video?
For explainers, tutorials, training and internal content it often is, provided the script is clean and someone reviews the result. For dramatic, comedic or emotionally loaded material, a human performer still does noticeably better. The comparison in AI dubbing vs human dubbing goes into where the line falls.