Modern Speech and Audio Architectures
Modern audio architectures evolved from dilated convolutions predicting individual sound samples to self-supervised transformers and neural codec language models. Whether converting text into natural voices, transcribing multilingual speech, or cloning voices with tokens, these models treat sound as structured sequences.
Why Does This Exist?
For decades, automated speech recognition (ASR) and text-to-speech (TTS) relied on complex, hand-engineered statistical pipelines. Speech recognition coupled acoustic hidden Markov models (HMMs), Gaussian mixture models (GMMs), pronunciation lexicons, and n-gram language models. Speech synthesis concatenated recorded diphones or ran statistical parametric vocoders that produced flat, buzzy, mechanical audio. Each submodule was trained in isolation under different objectives, accumulating compounding errors across pipeline boundaries.
The deep learning revolution transformed audio processing into end-to-end differentiable sequence modeling. Over the past decade, four milestone architectural paradigms emerged:
- WaveNet (2016): Proved neural networks can generate raw acoustic waveforms directly using causal dilated convolutions.
- Tacotron 2 (2018): Eliminated phoneme alignment pipelines by generating 80-bin Mel spectrograms directly from raw text characters via attention.
- Whisper (2022): Showed that scaling a standard encoder-decoder Transformer across 680,000 hours of weakly supervised audio matches human-level robustness across transcription, translation, and voice activity detection without task-specific heads.
- HuBERT & VALL-E (2021–2023): Discretized continuous audio into acoustic and semantic tokens via self-supervision and neural codecs (like EnCodec), framing speech synthesis as pure generative language modeling.
Think of It Like This
From hand-painted flipbooks to digital video tokens
Early speech systems were like artists manually assembling a flipbook: one specialist sketched phoneme outlines, another colored the pitch contours, and a third glued the pages together. If any hand slipped, the motion looked jagged and robotic.
WaveNet replaced the flipbook by calculating the exact shade of every microscopic pixel sequentially at 16,000 frames per second. Tacotron stepped back to render the high-level scene composition (the spectrogram) in one stroke. Whisper built an omniscient translator that watches any video stream, ignores background cafe noise, and outputs subtitles in 99 languages. Finally, VALL-E compressed the video into standardized Lego bricks (discrete acoustic tokens), allowing an ordinary text language model to clone anyone's voice from just three seconds of sample audio simply by continuing the sequence.
How It Actually Works
The Four Landmark Paradigms
1. WaveNet: Dilated Causal Convolutions
WaveNet models the joint probability of an audio waveform as a product of conditional probabilities:
To prevent future samples from leaking into the past, convolutions are strictly causal. Standard convolutions require thousands of layers to cover hundreds of milliseconds of audio. WaveNet introduces dilated convolutions, where filters skip input values with dilation factor . By doubling dilation exponentially across layers (), the receptive field grows exponentially with network depth while computational complexity scales linearly:
Each layer applies a gated activation unit modeled after PixelCNN:
where and denote filter and gate convolutions, and is element-wise multiplication.
2. Tacotron 2: Attention-Based Spectrogram Synthesis
Tacotron 2 maps text character sequences to Mel spectrogram frames. It consists of:
- Encoder: Character embeddings pass through a 3-layer convolutional filterbank followed by a Bidirectional LSTM to extract contextual phonemic representations.
- Location-Sensitive Attention: Computes alignment weights using both the decoder state and previous cumulative attention weights, preventing phoneme repetitions and omissions.
- Autoregressive Decoder: A 2-layer unidirectional LSTM that predicts one 80-bin Mel frame alongside a scalar stop token (trained via binary cross-entropy) indicating end-of-speech.
- Post-Net: A 5-layer convolutional network that predicts a residual detail tensor added to the decoder output, minimizing mean squared error before feeding into a neural vocoder.
3. Whisper: Weakly Supervised Multitask Transformer
OpenAI's Whisper formats all speech tasks (ASR, speech translation, voice activity detection, and word-level alignment) as sequence-to-sequence language modeling:
- Input Representation: 80-channel Log-Mel spectrogram computed with windows and hops.
- Stem: Two 1D convolutional layers with filter width 3 and stride 2 downsample the temporal dimension by a factor of 4.
- Encoder-Decoder Transformer: Standard sinusoidal position embeddings feed a Transformer encoder. The autoregressive decoder receives special prefix conditioning tokens:
Conditioning on tokens like switches the decoder from transcription to English translation without changing a single model parameter.
4. HuBERT & VALL-E: Discrete Acoustic Units and Codec LMs
- HuBERT (Hidden-Unit BERT): Takes continuous audio, applies span masking, and trains a BERT-like encoder to predict discrete cluster IDs obtained from -means clustering on intermediate feature representations. This forces the model to learn phonemic abstractions without human transcriptions.
- VALL-E: Uses a neural audio codec (EnCodec) that compresses audio into 8 hierarchical codebooks via Residual Vector Quantization (RVQ) at 75 tokens per second. An autoregressive language model generates the first codebook tokens conditioned on phoneme text and a 3-second acoustic prompt. Then, a non-autoregressive Transformer predicts codebooks through in parallel, achieving zero-shot voice cloning.
Worked Example
Let us compare the receptive field and computational scaling between WaveNet and Whisper.
-
WaveNet Exponential Receptive Field:
- Kernel size .
- Dilation cycle: 10 layers with .
- Receptive field per 10-layer block:
- Stacking 3 repeating cycles ( layers total):
- At a sampling rate ():
- Captures approximately two syllables of speech context without any pooling layers.
-
Whisper Conv1D Stem Temporal Downsampling:
- Input audio: at waveform samples.
- Mel spectrogram frames ( hop):
- Two Conv1D layers with stride 2 downsample time by :
- Full self-attention matrix comparisons drop from to operations—a computation reduction that makes 30-second context windows practical on standard GPUs.
Code
from typing import List
def wavenet_receptive_field(kernel_size: int, dilations: List[int], stacks: int) -> int: """Calculates the exact theoretical receptive field of a stacked WaveNet.""" single_stack_field = sum(d * (kernel_size - 1) for d in dilations) return 1 + stacks * single_stack_field
def whisper_sequence_length(duration_sec: float, hop_ms: float, strides: List[int]) -> int: """Calculates downsampled Transformer sequence length in Whisper.""" num_mel_frames = int((duration_sec * 1000.0) / hop_ms) total_downsample = 1 for s in strides: total_downsample *= s return num_mel_frames // total_downsample
# 1. Evaluate WaveNet with 3 stacks of 10 dilated layersdil_pattern = [2**i for i in range(10)] # [1, 2, 4, ..., 512]rf_samples = wavenet_receptive_field(kernel_size=2, dilations=dil_pattern, stacks=3)duration_ms = (rf_samples / 16000.0) * 1000.0
print(f"WaveNet Receptive Field: {rf_samples} samples ({duration_ms:.1f} ms)")# -> WaveNet Receptive Field: 3070 samples (191.9 ms)
# 2. Evaluate Whisper 30-second audio window downsamplingseq_len = whisper_sequence_length(duration_sec=30.0, hop_ms=10.0, strides=[2, 2])print(f"Whisper Transformer Sequence Length: {seq_len} tokens")# -> Whisper Transformer Sequence Length: 750 tokensWatch Out For
Autoregressive generation latency in sample-level vocoders
Symptom: A speech synthesis model takes 45 seconds of GPU computation to generate just 3 seconds of spoken audio, rendering it useless for interactive voice assistants.
Because WaveNet predicts raw audio sample-by-sample, generating 3 seconds of audio at requires sequential forward passes through the network. Even with optimized CUDA caching, sequential inference cannot saturate GPU tensor cores.
Modern production pipelines decouple semantic generation from audio rendering. Use FastSpeech 2 or VALL-E to generate Mel frames or discrete codec tokens in parallel, and synthesize raw waveforms using non-autoregressive GAN vocoders like HiFi-GAN or diffusion vocoders like DiffWave, which render entire audio clips in a single forward pass with real-time factor.
The Quick Version
- WaveNet proved that dilated causal convolutions can synthesize raw audio waveforms sample-by-sample without pooling.
- Tacotron 2 eliminated multi-stage TTS pipelines by using attention to map text directly to 80-channel Mel spectrograms.
- Whisper scales standard encoder-decoder Transformers across 680,000 hours, unifying speech tasks via special prefix tokens.
- HuBERT uses masked prediction of k-means cluster targets to learn rich phonetic representations without text transcriptions.
- VALL-E and modern codec models treat audio as discrete tokens, turning zero-shot speech generation into a language modeling problem.