Deep Time Series Architectures
Deep learning for time series replaces manual statistical modeling with neural building blocks designed specifically for temporal patterns. From interpretable backward-and-forward forecast stacks to causal dilated convolutions and patch-based transformers, these architectures capture complex multi-horizon dependencies.
Why Does This Exist?
Classical statistical forecasting methods like ARIMA, exponential smoothing (ETS), and vector autoregression (VAR) perform well on clean univariate series with short horizons. However, they struggle when scaling across thousands of interrelated series, long lookback windows (), non-linear dynamics, or multi-horizon forecasts ().
Early deep learning attempts simply applied standard recurrent networks (LSTMs and GRUs) to temporal sequences. But RNNs suffer from sequential compute bottlenecks, vanishing gradients over long horizons, and error compounding during iterative autoregressive rollouts.
To overcome these barriers, four landmark deep forecasting architectures emerged:
- N-BEATS (2019): Demonstrated that a pure multi-layer perceptron (MLP) with doubly residual connections can beat statistical ensembles on the M4 competition without recurrence or attention.
- Temporal Convolutional Networks (TCN, 2018): Replaced recurrent units with 1D causal dilated convolutions, achieving massive receptive fields and parallel training.
- Informer (2021): Solved long sequence time-series forecasting (LSTF) by replacing quadratic attention with ProbSparse attention.
- PatchTST (2023): Revolutionized Transformer-based forecasting through subseries patching and channel independence, establishing state-of-the-art accuracy while slashing memory footprint.
Think of It Like This
Four ways to predict the weather for the next month
Imagine an atmospheric research bureau testing four different predictive strategies:
- N-BEATS is the mathematical decomposition committee: It looks at historical barometric pressure, uses a dedicated desk to subtract the smooth multi-year warming trend, passes the remainder to another desk to subtract the repeating seasonal solstice wave, and sums their forward predictions.
- TCN is the telescope array: Instead of inspecting every second sequentially, it uses lenses spaced exponentially further apart () to see months into the past in parallel without losing temporal order.
- Informer is the executive summarizer: Rather than comparing every hour of last year with every other hour, it scans for anomalous storms, attends only to the most informative active weather queries, and predicts the entire month in one stroke.
- PatchTST is the subseries jigsaw specialist: Instead of treating single hourly numbers as tokens, it groups consecutive days into compact semantic patches and forecasts temperature, wind, and pressure through independent channels to keep sensor noise from bleeding across variables.
How It Actually Works
Architectural Mechanics and Formulations
1. N-BEATS: Doubly Residual Basis Expansions
N-BEATS operates as a cascade of blocks organized into stacks. Each block receives historical input (where is the original lookback window). The block consists of a 4-layer fully connected stack followed by two linear projection heads generating coefficient vectors (backcast) and (forecast):
These coefficients parameterize basis functions and :
- Backcast Output: reconstructs the historical input.
- Forecast Output: predicts the future horizon.
The block applies a double residual update:
- The backcast is subtracted from the block input to remove the signal it explained:
- The final forecast is the sum of all individual block forecasts across all stacks:
In interpretable N-BEATS, basis functions are constrained:
- Trend Basis: Monomial powers of time (linear, quadratic trends).
- Seasonality Basis: Periodic Fourier harmonics .
2. Temporal Convolutional Networks (TCN)
A TCN maps a sequence to using two fundamental principles:
- Causal Convolutions: A filter at timestep is strictly restricted to inputs at and earlier (). Convolutions are implemented with causal left-padding.
- Dilated Convolutions: A dilated convolution operation over sequence with filter is defined as: By increasing dilation exponentially with depth ( for layer ), the receptive field covers timesteps with network depth.
Residual blocks combine two dilated causal convolution layers, weight normalization, ReLU, spatial dropout, and a conv shortcut to prevent vanishing gradients.
3. Informer: ProbSparse Attention and Generative Decoder
Standard Transformer attention computes:
Evaluating all query-key dot products takes time and memory. Informer observes that the attention score distribution is heavy-tailed: only a small subset of "active" queries dominate.
Informer defines query sparsity measurement using Kullback-Leibler divergence against a uniform distribution:
ProbSparse Attention samples a subset of keys to compute , picks the top queries with the highest score, and fills non-selected query outputs with mean keys. This slashes attention complexity to .
In addition, Informer introduces distillation (halving sequence length between attention blocks via 1D convolutions with max pooling) and a generative direct multi-step decoder that avoids autoregressive drift.
4. PatchTST: Subseries Patching and Channel Independence
PatchTST achieved a major breakthrough in long-term forecasting by introducing two design paradigms:
- Patching: Instead of treating each individual timestep as a token, PatchTST segments univariate time series into overlapping patches of length with stride : Each patch is projected via a linear layer into a token embedding of dimension . Patching aggregates local temporal context, preserves semantic curves, and shrinks token sequence length from to , reducing self-attention compute by or .
- Channel Independence: In a multivariate series with channels, earlier models concatenated all variables into a single token vector. PatchTST splits the channels and treats each channel as an independent univariate time series that shares the exact same Transformer weights. This prevents inter-channel noise from corrupting temporal representations and yields significant empirical gains.
Worked Example
Let us quantify the computational savings of PatchTST over a standard pointwise Transformer on a typical long-term forecasting benchmark:
- Lookback window:
- Forecast horizon:
- Multivariate variables: (e.g., electricity load features)
- PatchTST configuration: Patch length , Stride
-
Calculate Pointwise Attention Matrix Size:
- Sequence length per channel:
- Self-attention matrix per head:
-
Calculate PatchTST Token Count per Channel:
- Number of patches :
-
Calculate PatchTST Attention Matrix Size:
- Self-attention matrix per head:
-
Compute Attention Memory Reduction Factor: Self-attention memory drops by over per channel.
-
Compare Total Operations Across All 7 Channels:
- Standard pointwise joint attention: attention entries.
- PatchTST channel-independent attention across all channels: Even while processing each channel independently, total attention compute is nearly smaller (), while preventing overfitting to spurious cross-sensor correlations.
Code
from typing import List, Tuple
def patchify_time_series( series: List[float], patch_len: int, stride: int) -> List[List[float]]: """Segments a 1D time series into overlapping subseries patches.""" patches = [] start = 0 while start + patch_len <= len(series): patch = series[start : start + patch_len] patches.append(patch) start += stride return patches
def compare_attention_complexity( seq_len: int, patch_len: int, stride: int, num_channels: int) -> Tuple[int, int, float]: """Compares attention matrix entries between pointwise and PatchTST architectures.""" pointwise_entries = (seq_len ** 2) num_patches = ((seq_len - patch_len) // stride) + 1 patchtst_entries_per_chan = (num_patches ** 2) total_patchtst_entries = patchtst_entries_per_chan * num_channels
reduction_factor = pointwise_entries / total_patchtst_entries return pointwise_entries, total_patchtst_entries, reduction_factor
# 1. Patch a sample time-series of 32 stepssynthetic_series = [float(i) for i in range(32)]patches_out = patchify_time_series(synthetic_series, patch_len=8, stride=4)
print(f"Original length: {len(synthetic_series)} steps -> Generated {len(patches_out)} patches:")for i, p in enumerate(patches_out[:3]): print(f"Patch {i}: {p}")# -> Original length: 32 steps -> Generated 7 patches:# -> Patch 0: [0.0, 1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0]# -> Patch 1: [4.0, 5.0, 6.0, 7.0, 8.0, 9.0, 10.0, 11.0]# -> Patch 2: [8.0, 9.0, 10.0, 11.0, 12.0, 13.0, 14.0, 15.0]
# 2. Benchmark attention matrix scaling (L=336, P=16, S=8, M=7)pw_cost, pt_cost, speedup = compare_attention_complexity( seq_len=336, patch_len=16, stride=8, num_channels=7)
print(f"Pointwise Attention Entries: {pw_cost:,}")# -> Pointwise Attention Entries: 112,896print(f"PatchTST Total Entries (7 channels): {pt_cost:,}")# -> PatchTST Total Entries (7 channels): 11,767print(f"Overall Attention Compute Reduction: {speedup:.1f}x")# -> Overall Attention Compute Reduction: 9.6xWatch Out For
Overfitting multivariate correlations with channel-dependent attention
Symptom: A Transformer model achieves lower training loss when concatenating multiple channels together, but achieves significantly worse test error than a simple linear model like DLinear.
In multivariate time series forecasting (such as predicting 300 stock tickers or 100 power grid buses), spurious temporal co-movements frequently arise. When a multi-head Transformer allows queries in channel to attend to keys in channel , the cross-channel attention heads easily memorize non-causal noise in the historical training window.
Unless physical cross-channel relationships are strictly verified and stationary (e.g., via domain graph adjacency matrices), default to Channel Independence. Treat each channel as an isolated univariate series processed by shared Transformer weights, as demonstrated by PatchTST.
The Quick Version
- N-BEATS uses doubly residual stacking of pure MLPs with interpretable polynomial and harmonic Fourier basis projections.
- Temporal Convolutional Networks (TCN) replace recurrence with causal dilated convolutions for parallelized training and large receptive fields.
- Informer introduces ProbSparse attention to cut long-horizon self-attention complexity from down to .
- PatchTST segments time series into overlapping subseries patches, extracting localized semantic patterns and reducing token counts.
- Channel independence processes multivariate variables as separate sequences with shared weights, eliminating cross-channel overfitting.