Audio and Speech Representations
Sound begins as rapid air pressure waves recorded thousands of times per second. By slicing these raw waveforms into short overlapping windows and calculating their frequency spectrum on a human-inspired Mel scale, models get clear visual-like audio representations.
Why Does This Exist?
Feeding raw 1D acoustic audio directly into deep neural networks presents severe statistical hurdles. At a standard speech sampling rate of , a single second of speech contains discrete scalar values. A short five-second sentence produces timesteps. Adjacent waveform samples are dominated by phase information and microscopic air-pressure oscillations that shift drastically if a speaker moves half an inch from the microphone, even though the phoneme spoken remains identical.
Directly processing raw audio in the time domain forces neural networks to expend massive capacity learning basic Fourier-like filters from scratch. Furthermore, human auditory perception does not perceive acoustic frequency linearly. The human cochlea resolves pitch with exquisite precision in low frequencies below (where speech formants distinguish vowels like "aa" vs "ee"), while grouping frequencies above into coarse perceptual bands.
Transforming 1D waveforms into 2D time-frequency representations—most notably the Short-Time Fourier Transform (STFT) and Log-Mel Spectrogram—compresses high-rate audio into compact matrices of approximately 100 frames per second, isolating invariant phonetic content from noisy acoustic phase.
Think of It Like This
A musical score transcribed from live violin vibrations
Imagine placing a high-speed camera in front of a violin string vibrating 440 times per second (). If you record the vertical position of the string at every microsecond, you get an overwhelming list of millions of raw numbers that obscure the tune.
Instead, a composer listens in short fractions of a second and marks which musical note is sounding on a sheet music staff. The vertical axis on the staff represents the pitch (arranged so that higher octaves space out naturally to the human ear), the horizontal axis represents time, and the darkness of the note represents loudness. A Mel spectrogram is simply automated sheet music for machines: it strips away the raw mechanical string flutter and leaves only the melody, harmonics, and rhythm.
How It Actually Works
From Continuous Waves to Log-Mel Spectrograms
The standard audio feature extraction pipeline proceeds through four distinct mathematical stages:
1. Sampling and Quantization (Nyquist-Shannon Theorem)
Continuous acoustic pressure is converted to a discrete sequence where . By the Nyquist-Shannon sampling theorem, capturing frequencies up to requires a sampling rate . For standard human speech (bandwidth up to ), is the universal standard. Samples are quantized into 16-bit signed integers (PCM format) spanning and normalized to .
2. Framing and Windowing
Speech signals are non-stationary over long intervals, but quasi-stationary over short windows of (the physical speed at which human vocal cords and articulators change shape). The signal is partitioned into overlapping frames of length ( at ) with hop size (). To prevent spectral leakage caused by abruptly truncating waveforms at frame boundaries, each frame is multiplied by a bell-shaped Hann window :
3. Short-Time Fourier Transform (STFT)
For each windowed frame , an -point Fast Fourier Transform (typically ) computes the discrete complex spectrum:
Because real-valued signals yield conjugate-symmetric spectra, only the first frequency bins are retained. The power spectrogram is the squared magnitude:
4. Mel Filterbank Integration and Log Compression
The linear frequency in Hertz is mapped to the perceptual Mel scale via:
A bank of overlapping triangular filters (typically for models like Whisper) spans from to . Each filter computes a weighted sum over the linear power bins:
Finally, human perception of loudness is logarithmic (measured in decibels). Applying a logarithm with numerical offset produces the final Log-Mel Spectrogram:
Worked Example
Let us trace the frequency-to-mel mapping and STFT bin assignment for a speech signal recorded at with bins and filters.
-
Calculate Linear Bin Spacing:
-
Locate a Fundamental Pitch:
-
Convert Linear Frequencies to Mel Scale:
- At (first vowel formant):
- At (second harmonic):
- At (third harmonic):
-
Observe Non-Linear Compression:
- Interval span:
- Interval span: The doubling from 800 to 1600 Hz covers only 1.38x the perceptual distance of 400 to 800 Hz, concentrating filterbank resolution where human phoneme distinctions occur.
Code
from typing import List, Tupleimport math
def hz_to_mel(hz: float) -> float: """Converts frequency in Hertz to perceptual Mel scale.""" return 2595.0 * math.log10(1.0 + hz / 700.0)
def mel_to_hz(mel: float) -> float: """Converts perceptual Mel scale back to Hertz.""" return 700.0 * (10.0 ** (mel / 2595.0) - 1.0)
def generate_mel_filterbank_centers( num_filters: int, f_min: float, f_max: float) -> List[float]: """Computes peak frequencies (Hz) for linearly spaced Mel filters.""" mel_min = hz_to_mel(f_min) mel_max = hz_to_mel(f_max) step = (mel_max - mel_min) / (num_filters + 1) mel_points = [mel_min + i * step for i in range(1, num_filters + 1)] return [mel_to_hz(m) for m in mel_points]
# Compute filter center frequencies for standard 16 kHz audio setupcenters_hz = generate_mel_filterbank_centers(num_filters=5, f_min=0.0, f_max=8000.0)
print("Mel filter center frequencies (Hz):")for idx, freq in enumerate(centers_hz, 1): print(f"Filter {idx}: {freq:.1f} Hz")# -> Mel filter center frequencies (Hz):# -> Filter 1: 342.3 Hz# -> Filter 2: 875.9 Hz# -> Filter 3: 1707.9 Hz# -> Filter 4: 3005.1 Hz# -> Filter 5: 5027.6 Hz
# Notice how spacing widens:diffs = [centers_hz[i+1] - centers_hz[i] for i in range(len(centers_hz)-1)]print("Bandwidth gaps:", [round(d, 1) for d in diffs])# -> Bandwidth gaps: [533.6, 832.0, 1297.2, 2022.5]Watch Out For
Phase erasure during spectrogram synthesis without neural vocoders
Symptom: Audio generated by inverting spectrograms with Griffin-Lim sounds robotic, hollow, or metallic with severe background phase buzzing.
The short-time Fourier transform produces complex values . Converting to power and Mel filterbanks discards the phase completely. The classical Griffin-Lim algorithm iteratively estimates missing phase through repeated forward and inverse STFTs, but it converges slowly and generates distinct metallic artifacts.
In modern generative pipelines (such as text-to-speech or audio super-resolution), never rely on naive phase reconstruction algorithms. Always pair Mel spectrogram generation with a trained neural vocoder (such as HiFi-GAN, WaveNet, or BigVGAN) that learns the non-linear prior distribution of phase directly from data.
The Quick Version
- Raw acoustic audio at has high redundancy, noise sensitivity, and massive temporal length ( points/sec).
- Short-Time Fourier Transform (STFT) windows raw audio into frames to extract localized frequency spectra.
- The Mel scale aligns frequency bin resolution with human ear cochlear sensitivity, emphasizing lower speech formants.
- Overlapping triangular Mel filters followed by logarithmic power compression produce the standard 80-bin Log-Mel spectrogram.
- Log-Mel representations are standard inputs for modern speech recognition models like Whisper and speech synthesizers like Tacotron.