GenLucid

SPEECH SYNTHESIS

How Does Text-to-Speech Work?

Neural text-to-speech runs in stages. Text is converted into a sequence of sounds (phonemes), an acoustic model predicts a spectrogram that describes what the audio should look like over time, and a vocoder turns that spectrogram into the waveform a speaker plays. Pitch and timing are predicted along the way, and they are what make the same words sound like a statement, a question, or an exclamation.

One-pager titled "How does text-to-speech work?": the word "hello" becomes four sound chips (HH, AH, L, OW), an acoustic model turns them into a spectrogram where each sound's column width is how long it lasts and a cream line traces the pitch, and a vocoder turns the spectrogram into a waveform; below, the same word is shown three ways, "Hello." with a falling pitch line, "Hello?" with a rising line, and "Hello!" higher and faster; a REMEMBER band notes letters become sounds first, pitch plus timing equals delivery, and a vocoder makes the wave.
One-pager: text becomes sounds, an acoustic model predicts a spectrogram, and a vocoder renders the waveform, with pitch and timing setting the delivery.

Step 1: text becomes a sequence of sounds

Before any audio exists, the system has to decide what to say out loud. Written text is not a pronunciation guide. "3" has to become "three", "Dr." can be "doctor" or "drive", and English spelling gives no reliable clue to how "through" and "though" sound. Many systems first normalize the text into plain words, then convert the words into phonemes, the distinct sound units of speech. On the one-pager, "hello" becomes four sounds: HH, AH, L, OW.

This front end is a design choice, not a law. FastSpeech 2 takes a phoneme sequence as its input. Tacotron 2 takes character embeddings and learns pronunciation inside the network instead.

Step 2: an acoustic model predicts a spectrogram

The acoustic model says what the sound should look like over time. Its output is a mel spectrogram: time runs left to right, pitch runs bottom to top on a scale spaced the way human hearing is, and brightness is how much energy sits at that pitch at that moment. In FastSpeech 2's setup, each column advances by 256 samples of 22,050 Hz audio, about 12 milliseconds, and holds 80 values.

The picture is readable. A vowel such as AH shows up as bright horizontal bands, because a vowel holds steady for a while. A breathy consonant such as HH shows up as a short smear of energy spread across many pitches. Sounds that last longer take up more columns, which is why the one-pager draws OW wider than L.

Two families of acoustic model exist. Tacotron 2 is a recurrent sequence-to-sequence network that maps character embeddings to mel spectrograms one step after another. FastSpeech 2 is non-autoregressive: it processes the whole phoneme sequence in parallel and outputs the whole spectrogram at once.

Step 3: duration, pitch and energy set the delivery

Words alone do not fix how a sentence is spoken. FastSpeech 2 adds a variance adaptor that predicts three things: how long each phoneme lasts, and the pitch and energy of the speech. Those predictions are added into the model's hidden sequence before the spectrogram is produced.

During training, the duration targets come from a forced aligner (the Montreal Forced Aligner in the paper), a tool that finds where each phoneme starts and ends in a real recording. The model learns its patterns from many recordings. A habit such as raising pitch at the end of a question is learned from data rather than coded as a rule.

The bottom row of the one-pager illustrates the effect. The same word gets a falling pitch line as a statement, a rising line as a question, and a higher, faster line when excited. The phonemes are identical in all three. Only the pitch curve and the timing change.

Step 4: a vocoder turns the spectrogram into a waveform

A spectrogram is a plan, not audio. A speaker needs a waveform, which is tens of thousands of values per second, and the vocoder generates them from the spectrogram.

WaveNet generates audio one sample at a time, with each sample conditioned on all the previous ones. In the paper, human listeners rated it as significantly more natural than the best parametric and concatenative systems, for both English and Mandarin. Tacotron 2 paired a modified WaveNet with its spectrogram predictor and reported a mean opinion score of 4.53, against 4.58 for professionally recorded speech. Its ablations found that using mel spectrograms as the WaveNet input allowed a significantly simpler WaveNet.

Sample-by-sample generation is slow. HiFi-GAN trains a generative adversarial network to produce the raw waveform directly, and its authors report 22.05 kHz audio generated 167.9 times faster than real time on a single V100 GPU. The design idea they stress is modeling the periodic patterns in audio, since speech consists of sinusoidal signals with various periods.

Where the stages blur together

The four-stage picture is the clearest way to see the mechanism, not the only architecture. The HiFi-GAN authors also show their vocoder inside an end-to-end speech synthesis setup, and some newer speech models predict discrete audio codec tokens instead of a spectrogram, then decode those tokens into a waveform (see what a neural audio codec is). What stays constant is the split between deciding what the sound should be and rendering that decision as audio.

FAQ

How does text-to-speech read numbers and abbreviations?
Most pipelines run a text normalization step first that rewrites them as words, so "3 cats" becomes "three cats" before pronunciation is predicted. Ambiguous cases, such as an abbreviation with two possible readings, are where errors tend to appear.
Why does AI text-to-speech mispronounce names?
Pronunciation is predicted from patterns in training data. A rare name has few examples, so the predicted phoneme sequence can be wrong, and the acoustic model then renders those wrong sounds faithfully. Some systems let you type the phonemes in directly, which skips the prediction.
How is neural text-to-speech different from older robotic voices?
Concatenative systems join fragments of recorded speech, and parametric systems generate audio from hand-designed acoustic models. Neural systems predict the audio itself. In the WaveNet paper, human listeners rated its output as significantly more natural than the best parametric and concatenative systems of the time.
Can text-to-speech copy a specific voice?
Yes, when the acoustic model is also given a speaker embedding, a vector describing a voice, extracted from a short reference clip. See how AI voice cloning works for the details.

Sources

Related

Last updated 2026-09-29