Skip to content
AI360Xpert
Beta

Audio and Speech Representations

Sound begins as rapid air pressure waves recorded thousands of times per second. By slicing these raw waveforms into short overlapping windows and calculating their frequency spectrum on a human-inspired Mel scale, models get clear visual-like audio representations.

End-to-end audio processing pipeline transforming continuous waveform to PCM samples, STFT spectrogram, and 80-channel log-Mel filterbank
End-to-end audio processing pipeline transforming continuous waveform to PCM samples, STFT spectrogram, and 80-channel log-Mel filterbank

Why Does This Exist?

Feeding raw 1D acoustic audio directly into deep neural networks presents severe statistical hurdles. At a standard speech sampling rate of 16 kHz16\text{ kHz}, a single second of speech contains 16,00016{,}000 discrete scalar values. A short five-second sentence produces 80,00080{,}000 timesteps. Adjacent waveform samples are dominated by phase information and microscopic air-pressure oscillations that shift drastically if a speaker moves half an inch from the microphone, even though the phoneme spoken remains identical.

Directly processing raw audio in the time domain forces neural networks to expend massive capacity learning basic Fourier-like filters from scratch. Furthermore, human auditory perception does not perceive acoustic frequency linearly. The human cochlea resolves pitch with exquisite precision in low frequencies below 1 kHz1\text{ kHz} (where speech formants distinguish vowels like "aa" vs "ee"), while grouping frequencies above 4 kHz4\text{ kHz} into coarse perceptual bands.

Transforming 1D waveforms into 2D time-frequency representations—most notably the Short-Time Fourier Transform (STFT) and Log-Mel Spectrogram—compresses high-rate audio into compact matrices of approximately 100 frames per second, isolating invariant phonetic content from noisy acoustic phase.

Think of It Like This

A musical score transcribed from live violin vibrations

Imagine placing a high-speed camera in front of a violin string vibrating 440 times per second (440 Hz440\text{ Hz}). If you record the vertical position of the string at every microsecond, you get an overwhelming list of millions of raw numbers that obscure the tune.

Instead, a composer listens in short fractions of a second and marks which musical note is sounding on a sheet music staff. The vertical axis on the staff represents the pitch (arranged so that higher octaves space out naturally to the human ear), the horizontal axis represents time, and the darkness of the note represents loudness. A Mel spectrogram is simply automated sheet music for machines: it strips away the raw mechanical string flutter and leaves only the melody, harmonics, and rhythm.

How It Actually Works

From Continuous Waves to Log-Mel Spectrograms

The standard audio feature extraction pipeline proceeds through four distinct mathematical stages:

1. Sampling and Quantization (Nyquist-Shannon Theorem)

Continuous acoustic pressure x(t)x(t) is converted to a discrete sequence x[n]=x(n⋅Ts)x[n] = x(n \cdot T_s) where Ts=1/fsT_s = 1/f_s. By the Nyquist-Shannon sampling theorem, capturing frequencies up to fmaxf_{\text{max}} requires a sampling rate fs≥2fmaxf_s \ge 2 f_{\text{max}}. For standard human speech (bandwidth up to 8 kHz8\text{ kHz}), fs=16,000 Hzf_s = 16{,}000\text{ Hz} is the universal standard. Samples are quantized into 16-bit signed integers (PCM format) spanning [−32768,32767][-32768, 32767] and normalized to [−1.0,1.0][-1.0, 1.0].

2. Framing and Windowing

Speech signals are non-stationary over long intervals, but quasi-stationary over short windows of 20 to 30 ms20\text{ to }30\text{ ms} (the physical speed at which human vocal cords and articulators change shape). The signal is partitioned into overlapping frames of length N=400N = 400 (25 ms25\text{ ms} at 16 kHz16\text{ kHz}) with hop size H=160H = 160 (10 ms10\text{ ms}). To prevent spectral leakage caused by abruptly truncating waveforms at frame boundaries, each frame is multiplied by a bell-shaped Hann window w[n]w[n]:

w[n]=0.5−0.5cos⁡(2πnN−1),0≤n≤N−1w[n] = 0.5 - 0.5 \cos\left( \frac{2\pi n}{N - 1} \right), \quad 0 \le n \le N-1

3. Short-Time Fourier Transform (STFT)

For each windowed frame mm, an NfftN_{\text{fft}}-point Fast Fourier Transform (typically Nfft=512N_{\text{fft}} = 512) computes the discrete complex spectrum:

X[m,k]=∑n=0N−1x[m⋅H+n]⋅w[n]⋅e−j2πknNfft,0≤k≤Nfft2X[m, k] = \sum_{n=0}^{N-1} x[m \cdot H + n] \cdot w[n] \cdot e^{-j \frac{2\pi k n}{N_{\text{fft}}}}, \quad 0 \le k \le \frac{N_{\text{fft}}}{2}

Because real-valued signals yield conjugate-symmetric spectra, only the first Nfft/2+1=257N_{\text{fft}}/2 + 1 = 257 frequency bins are retained. The power spectrogram is the squared magnitude:

P[m,k]=∣X[m,k]∣2=Re⁡(X[m,k])2+Im⁡(X[m,k])2P[m, k] = |X[m, k]|^2 = \operatorname{Re}(X[m, k])^2 + \operatorname{Im}(X[m, k])^2

4. Mel Filterbank Integration and Log Compression

The linear frequency ff in Hertz is mapped to the perceptual Mel scale mm via:

m=2595log⁡10(1+f700)m = 2595 \log_{10}\left(1 + \frac{f}{700}\right)

A bank of MM overlapping triangular filters (typically M=80M = 80 for models like Whisper) spans from 0 Hz0\text{ Hz} to fs/2=8000 Hzf_s/2 = 8000\text{ Hz}. Each filter Hi[k]H_i[k] computes a weighted sum over the linear power bins:

Smel[m,i]=∑k=0Nfft/2P[m,k]⋅Hi[k],1≤i≤MS_{\text{mel}}[m, i] = \sum_{k=0}^{N_{\text{fft}}/2} P[m, k] \cdot H_i[k], \quad 1 \le i \le M

Finally, human perception of loudness is logarithmic (measured in decibels). Applying a logarithm with numerical offset ϵ\epsilon produces the final Log-Mel Spectrogram:

Y[m,i]=log⁡(Smel[m,i]+ϵ)Y[m, i] = \log(S_{\text{mel}}[m, i] + \epsilon)

Worked Example

Let us trace the frequency-to-mel mapping and STFT bin assignment for a speech signal recorded at fs=16,000 Hzf_s = 16{,}000\text{ Hz} with Nfft=512N_{\text{fft}} = 512 bins and M=80M = 80 filters.

  1. Calculate Linear Bin Spacing: Δf=fsNfft=16000512=31.25 Hz per bin\Delta f = \frac{f_s}{N_{\text{fft}}} = \frac{16000}{512} = 31.25\text{ Hz per bin}

  2. Locate a 400 Hz400\text{ Hz} Fundamental Pitch: k=400Δf=40031.25=12.8  ⟹  Bin index 13k = \frac{400}{\Delta f} = \frac{400}{31.25} = 12.8 \implies \text{Bin index } 13

  3. Convert Linear Frequencies to Mel Scale:

    • At f=400 Hzf = 400\text{ Hz} (first vowel formant): m(400)=2595log⁡10(1+400700)=2595log⁡10(1.5714)=2595×0.1963=509.4 melsm(400) = 2595 \log_{10}\left(1 + \frac{400}{700}\right) = 2595 \log_{10}(1.5714) = 2595 \times 0.1963 = 509.4\text{ mels}
    • At f=800 Hzf = 800\text{ Hz} (second harmonic): m(800)=2595log⁡10(1+800700)=2595log⁡10(2.1429)=2595×0.3310=859.0 melsm(800) = 2595 \log_{10}\left(1 + \frac{800}{700}\right) = 2595 \log_{10}(2.1429) = 2595 \times 0.3310 = 859.0\text{ mels}
    • At f=1600 Hzf = 1600\text{ Hz} (third harmonic): m(1600)=2595log⁡10(1+1600700)=2595log⁡10(3.2857)=2595×0.5166=1340.6 melsm(1600) = 2595 \log_{10}\left(1 + \frac{1600}{700}\right) = 2595 \log_{10}(3.2857) = 2595 \times 0.5166 = 1340.6\text{ mels}
  4. Observe Non-Linear Compression:

    • Interval 400 Hz→800 Hz400\text{ Hz} \to 800\text{ Hz} span: 859.0−509.4=349.6 mels859.0 - 509.4 = 349.6\text{ mels}
    • Interval 800 Hz→1600 Hz800\text{ Hz} \to 1600\text{ Hz} span: 1340.6−859.0=481.6 mels1340.6 - 859.0 = 481.6\text{ mels} The doubling from 800 to 1600 Hz covers only 1.38x the perceptual distance of 400 to 800 Hz, concentrating filterbank resolution where human phoneme distinctions occur.

Code

from typing import List, Tupleimport math
def hz_to_mel(hz: float) -> float:    """Converts frequency in Hertz to perceptual Mel scale."""    return 2595.0 * math.log10(1.0 + hz / 700.0)

def mel_to_hz(mel: float) -> float:    """Converts perceptual Mel scale back to Hertz."""    return 700.0 * (10.0 ** (mel / 2595.0) - 1.0)

def generate_mel_filterbank_centers(    num_filters: int,     f_min: float,     f_max: float) -> List[float]:    """Computes peak frequencies (Hz) for linearly spaced Mel filters."""    mel_min = hz_to_mel(f_min)    mel_max = hz_to_mel(f_max)    step = (mel_max - mel_min) / (num_filters + 1)    mel_points = [mel_min + i * step for i in range(1, num_filters + 1)]    return [mel_to_hz(m) for m in mel_points]

# Compute filter center frequencies for standard 16 kHz audio setupcenters_hz = generate_mel_filterbank_centers(num_filters=5, f_min=0.0, f_max=8000.0)
print("Mel filter center frequencies (Hz):")for idx, freq in enumerate(centers_hz, 1):    print(f"Filter {idx}: {freq:.1f} Hz")# -> Mel filter center frequencies (Hz):# -> Filter 1: 342.3 Hz# -> Filter 2: 875.9 Hz# -> Filter 3: 1707.9 Hz# -> Filter 4: 3005.1 Hz# -> Filter 5: 5027.6 Hz
# Notice how spacing widens:diffs = [centers_hz[i+1] - centers_hz[i] for i in range(len(centers_hz)-1)]print("Bandwidth gaps:", [round(d, 1) for d in diffs])# -> Bandwidth gaps: [533.6, 832.0, 1297.2, 2022.5]

Watch Out For

Phase erasure during spectrogram synthesis without neural vocoders

Symptom: Audio generated by inverting spectrograms with Griffin-Lim sounds robotic, hollow, or metallic with severe background phase buzzing.

The short-time Fourier transform produces complex values X[m,k]=∣X∣ejϕX[m, k] = |X| e^{j \phi}. Converting to power and Mel filterbanks discards the phase ϕ\phi completely. The classical Griffin-Lim algorithm iteratively estimates missing phase through repeated forward and inverse STFTs, but it converges slowly and generates distinct metallic artifacts.

In modern generative pipelines (such as text-to-speech or audio super-resolution), never rely on naive phase reconstruction algorithms. Always pair Mel spectrogram generation with a trained neural vocoder (such as HiFi-GAN, WaveNet, or BigVGAN) that learns the non-linear prior distribution of phase directly from data.

The Quick Version

  • Raw acoustic audio at 16 kHz16\text{ kHz} has high redundancy, noise sensitivity, and massive temporal length (16,00016{,}000 points/sec).
  • Short-Time Fourier Transform (STFT) windows raw audio into 25 ms25\text{ ms} frames to extract localized frequency spectra.
  • The Mel scale aligns frequency bin resolution with human ear cochlear sensitivity, emphasizing lower speech formants.
  • Overlapping triangular Mel filters followed by logarithmic power compression produce the standard 80-bin Log-Mel spectrogram.
  • Log-Mel representations are standard inputs for modern speech recognition models like Whisper and speech synthesizers like Tacotron.