AI voice & dubbing

Inside a Synthetic Voice: The Neural TTS Pipeline Step by Step

AI voice generation works by turning text into a sequence of linguistic units, predicting an intermediate representation of the sound (a spectrogram or a stream of discrete speech tokens), and converting that into a waveform with a neural vocoder or decoder. A separate conditioning signal, often taken from a short reference recording, tells the model whose voice to imitate. Each stage has its own failure modes, which explains most of the quirks you hear in synthetic speech.

7 min read · Updated

The pipeline at a glance

Whatever the branding, almost every modern neural text-to-speech system can be described as four stages. Some systems fuse two or three of them into a single network, but the jobs remain.

  1. Front end: clean up and interpret the text, then convert it into units the model understands.
  2. Acoustic modeling: predict what the speech should sound like over time, as a spectrogram or as discrete audio tokens.
  3. Waveform generation: turn that representation into actual audio samples.
  4. Conditioning: throughout, steer the output toward a particular speaker, language, style or emotional intensity.

The rest of this article goes through each stage, then looks at how a model can imitate a voice from a few seconds of audio. For the broader history of the field, see what is text to speech.

Text normalization and the front end

Raw text is full of things nobody pronounces as written. Normalization rewrites them into speakable words: "$4.99" becomes "four dollars ninety-nine", "3/4" becomes "three quarters" or "March fourth" or "the third of April" depending on context and locale, "St." becomes "street" or "saint". Rule-based normalizers handle the predictable cases; learned models increasingly handle the ambiguous ones from context.

Next comes grapheme-to-phoneme conversion: mapping spelling to sounds. English is notoriously irregular, so traditional systems used a pronunciation dictionary plus a learned model for unknown words. Many end-to-end neural models skip explicit phonemes and learn pronunciation straight from characters or subword tokens, which works well for common words and less well for rare names. Multilingual models must also know which language a word belongs to, since the same spelling is pronounced differently across languages.

Normalization in action

Input: "Dr. Lee moved to 12 Oak Dr. on 5/6 and paid $1.2M." A naive reading says "D R Lee moved to twelve Oak D R on five slash six." A good front end says "Doctor Lee moved to twelve Oak Drive on May sixth (or the fifth of June, in British usage) and paid one point two million dollars." The two abbreviations are spelled the same and pronounced differently, and the date depends on locale, which is why scripts meant for synthesis should be written out unambiguously.

Acoustic models: predicting the shape of the sound

The first wave of neural TTS predicted a mel spectrogram: a picture of how energy is spread across frequencies over time, warped to match human hearing. Tacotron and Tacotron 2 did this autoregressively, generating one frame after another while an attention mechanism decided which part of the text to read next. The results were natural, but attention could slip, causing skipped or repeated words.

Non-autoregressive models such as FastSpeech fixed the alignment problem with an explicit duration predictor: the model decides how many frames each sound should last, then generates all frames in parallel. That makes synthesis faster and more robust and also gives a direct handle on speaking rate. Models like VITS went further and trained the acoustic model and waveform generator jointly.

Speech-token language models

A newer family treats speech the way large language models treat text. A neural audio codec compresses audio into a sequence of discrete tokens, a few dozen to a few hundred per second. A transformer is then trained to predict those tokens from text, exactly as a text model predicts the next word. Microsoft's VALL-E (2023) popularized the idea, and many open models since follow a similar pattern: a language model produces coarse speech tokens, and a second network, often based on diffusion or flow matching, turns them into detailed acoustic features.

The appeal is that these models learn from very large, varied speech collections and pick up prosody, accent and recording character as part of the token sequence. The cost is that generation is sampled rather than fixed, so the same sentence can come out slightly differently each time, and a sampling misstep can produce a garbled syllable or an unexpected pause.

Vocoders: turning representations into audio

A spectrogram is not sound; it lacks phase information and fine detail. The vocoder fills that gap. WaveNet showed in 2016 that a neural network could generate raw audio sample by sample with remarkable quality, though slowly. Later vocoders such as WaveRNN, WaveGlow and the GAN-based HiFi-GAN made the same job fast enough for real-time use.

In token-based systems, the codec's own decoder or a dedicated vocoder plays this role. Vocoder quality is where you hear metallic edges, hiss or a slightly phasey sound, and it is also why a voice that is fine on clean text can degrade when the acoustic model hands over an unusual spectrogram.

Conditioning a voice on a short reference recording

To speak in a particular voice, the model needs a description of that voice. Older multi-speaker systems learned a fixed embedding per training speaker, so only voices in the training set were available. Zero-shot systems instead compute a representation from a reference clip at generation time. That might be a speaker embedding from a separate speaker-verification network, or, in token models, the reference audio's tokens placed at the start of the sequence as a prompt the model continues in the same voice.

Reference conditioning copies more than timbre. Room acoustics, microphone color, background hum and speaking style in the sample all leak into the output, so a clean, steady reference produces a clean, steady voice. It also raises obvious consent questions, since a few seconds of anyone's speech can be enough to imitate them convincingly.

What the neural pipeline does better than older systems

Splitting the job this way, and learning each stage from data, is what lifted synthetic speech from tolerable to pleasant. Learned prosody replaces hand-tuned intonation rules, so sentences sound shaped rather than recited. Neural vocoders removed the buzz of parametric systems and the audible joins of concatenative ones. Reference conditioning means a new voice no longer requires a voice actor to record for days. And because one model can cover many languages, the same voice can follow content across a translation workflow.

Limitations of each stage: where the pipeline breaks

Each stage contributes its own characteristic errors, and recognizing them helps you diagnose a bad result.

  • Front end errors: wrong expansion of numbers and abbreviations, wrong pronunciation of heteronyms and names.
  • Alignment errors: skipped words, repeated phrases or runaway babble, mostly in autoregressive models and on unusually long inputs.
  • Prosody errors: flat delivery, emphasis on the wrong word, unnatural pauses, especially on sentence fragments that lack context.
  • Vocoder artifacts: buzz, metallic ring or crackle, often on very high or very low pitches.
  • Conditioning leakage: echo or noise copied from a poor reference recording, or accent bleeding from the reference language into another language.
  • Sampling variance: a sentence that sounds fine on one generation and odd on the next.

How mydubly's voice stage fits this architecture

mydubly's default voice engine is Chatterbox Multilingual, an open-source model from Resemble AI; the Chatterbox TTS explainer covers what it offers. Some of mydubly's preset voices are produced by conditioning the model on a reference recording, in the way described above, and you choose one of 8 stock voices; you cannot upload your own voice.

Two engineering choices respond directly to the failure modes in this article. Short transcript fragments are merged into whole sentences of up to about 220 characters before synthesis, because the model produces steadier prosody on complete sentences. Afterwards each clip is normalized to mono 24 kHz with edge silence trimmed and consistent loudness, then placed against the original timing with at most a modest tempo change before the clips are joined into one AAC track. The details of that timing step are in syncing translated audio with video, and the overall voice product is on the AI dubbing page.

Where to go next

Understanding the pipeline makes you a sharper listener: you can tell a front end mistake from a vocoder artifact and fix the script accordingly. To hear the whole chain on real speech, try a short clip with AI dubbing; the minimum charge is 2 minutes, which is 100 credits or $0.10. If you are judging the output, our notes on AI voice quality explain how listening tests are run.

Frequently asked questions

What is the difference between an acoustic model and a vocoder?

The acoustic model decides what the speech should contain over time, such as pitch, duration and spectral shape, usually as a mel spectrogram or speech tokens. The vocoder converts that representation into actual audio samples you can play. Some modern systems train both together or merge them into one network.

Why does the same sentence sometimes sound different each time it is generated?

Many modern models sample from a probability distribution rather than computing one fixed answer, especially token-based models. Small random choices change timing, emphasis and intonation. Some tools fix the random seed for repeatable output, at the cost of not being able to re-roll a bad take.

How much audio is needed to imitate a voice?

Zero-shot systems can work from a few seconds of reference audio, though longer, clean samples usually capture the voice more faithfully. Fine-tuning a model on a specific speaker typically uses much more recorded material. In every case, a quiet, single-speaker recording matters more than length.

Do AI voices understand what they are saying?

Not in a human sense. The model learns statistical associations between text and sound, including that questions tend to rise and lists have rhythm. It does not know which word is important to the speaker's argument, which is why emphasis in synthetic speech can land in the wrong place.

Why do synthetic voices struggle with very short phrases?

A two-word fragment gives the model almost no context to infer intonation, so it often produces a flat or oddly final-sounding reading. Longer, complete sentences give it the grammatical cues it needs. That is the reason some pipelines merge fragments into full sentences before synthesis.