SPEECH RECOGNITION
How Does Speech-to-Text Work?
Speech-to-text, also called automatic speech recognition, turns audio into a spectrogram, encodes that spectrogram into features that describe the sound, and decodes the features into words one token at a time. The decoder also reads the words it has already written, which is how the same sound can become "there" in one sentence and "their" in another.
Step 1: sound is sliced into frames and drawn as a spectrogram
A microphone produces a waveform, a long list of pressure readings, and in that form words are hard for a model to see. Whisper first resamples all audio to 16,000 Hz, then computes an 80-channel log-magnitude Mel spectrogram on 25-millisecond windows with a stride of 10 milliseconds. That gives one new spectrogram column every 10 milliseconds. Whisper works on 30-second chunks, so one chunk becomes about 3,000 columns of 80 values.
It is the same kind of picture a text-to-speech system predicts (see how text-to-speech works), used here as the input instead of the output. Vowels appear as bright bands, and the shape of the energy over time is the evidence for which sounds were spoken.
Step 2: an encoder turns the spectrogram into features
Whisper's encoder starts with a stem of two convolution layers (filter width 3, GELU activation), the second with a stride of two. Sinusoidal position embeddings are added, and transformer blocks are applied. The convolutions pick up local patterns such as individual sounds. The stride of two halves the number of positions, so a 30-second chunk becomes about 1,500 positions. The transformer layers then relate every position to every other one (see what attention is), which is how sounds combine into words and words into phrases.
The output is one feature vector per position. These are not words yet. Each vector describes the sound around that moment in a form the decoder can read.
An encoder can also be pretrained without transcripts. wav2vec 2.0 masks parts of the speech in latent space and solves a contrastive task over quantized latent representations, so it learns from audio alone, then it is fine-tuned on transcribed speech. Its authors report word error rates of 4.8 and 8.2 on LibriSpeech test sets using only ten minutes of labeled data plus 53,000 hours of unlabeled audio.
Step 3: a decoder writes the transcript one token at a time
The decoder is a transformer that generates text the way a text LLM does: one token, then the next, each conditioned on the encoder's features and on every token written so far (see how next-token prediction works). At each step it scores every candidate token and picks one, and the pick is appended to the sequence.
Whisper steers this single decoder with special tokens: a language token, a transcribe or translate task token, a token that says whether to predict timestamps, and a token for segments with no speech. Timestamps are predicted as tokens too. Times are relative to the current audio segment, quantized to the nearest 20 milliseconds, and interleaved with the caption tokens.
The training set was 680,000 hours of labeled audio. Of that, 117,000 hours cover 96 languages other than English, and 125,000 hours are translation data from other languages into English.
Why the words already written matter
Sound does not always decide the spelling. "There", "their" and "they're" sound the same, so the features alone leave all three about equally likely, as the left panel of the one-pager draws it. The decoder also sees the words so far. After "put it over", "there" fits and the other two do not, so its score rises far above theirs. The bar lengths on the one-pager illustrate the idea and are not measured values.
The same mechanism explains some failures. With little context and few similar training examples, a common word that sounds close can outscore a rare name.
The other design: CTC labels every frame
Not every recognizer uses a text-generating decoder. Connectionist temporal classification (CTC) has the network output one label per frame, either a character or a special blank. Repeated labels are merged, then blanks are dropped. A blank must sit between two identical letters to keep both, so the frame sequence h, h, e, blank, l, l, blank, l, o becomes "hello".
CTC does not need to be told where each word starts in the audio, which suits training on audio paired with plain transcripts. Its cost, as Awni Hannun describes it in Distill, is the assumption that every output is conditionally independent of the other outputs given the input. A separate language model can be added on top, and it usually improves accuracy.
How accuracy is measured
Speech-to-text is scored with word error rate. Count the substitutions, deletions and insertions needed to turn the model's output into a reference transcript, then divide by the number of words in the reference. If the reference is "the cat sat" and the output is "the cat sat down", that is one insertion over three words, a word error rate of about 33 percent. Lower is better.
FAQ
- Can speech-to-text tell who is speaking?
- Not by itself. Base models such as Whisper transcribe what was said. Working out who said it, called speaker diarization, is a separate task usually handled by another model or processing step.
- Can speech-to-text translate as well as transcribe?
- Whisper can. A task token tells the decoder to translate into English instead of transcribing, so one model does both in a single pass. Its training data included 125,000 hours of translation into English.
- Why does speech-to-text struggle with accents and background noise?
- The model maps sounds to text using patterns from the audio it was trained on, so conditions that were thin in that data are handled less reliably. Whisper's authors trained on a large and varied set of labeled audio so that recognition holds up across more conditions.
- How does speech-to-text handle long recordings?
- Whisper processes audio in 30-second chunks, so a long recording is split and the chunks are decoded one after another, with timestamps given relative to each segment.
Sources
Related
Last updated 2026-09-29