Time Series and Forecasting Fundamentals
Time series data records measurements ordered sequentially across time, where the future depends on patterns from the past. Building reliable forecasts requires teasing apart underlying trends, repeating cycles, and noise while strictly preventing future data from leaking into past training steps.
Why Does This Exist?
Standard supervised machine learning assumes observations are independently and identically distributed (). In tabular classification or image processing, shuffling rows or swapping dataset order has zero impact on the mathematical validity of the model.
Time series data violently breaks the assumption. In sequential temporal measurements—such as electricity grid load, server CPU utilization, financial asset prices, or retail product demand—an observation at time is strongly correlated with observations at , and . Applying standard machine learning techniques out of the box leads to catastrophic pitfalls:
- Temporal Data Leakage: Running random -fold cross-validation trains the model on future observations (e.g., ) to predict past observations (e.g., ). This produces deceptively stellar validation scores that instantly collapse when deployed to real production environments.
- Spurious Regressions from Non-Stationarity: Regressing two unrelated time series that both exhibit an upward secular trend (such as global average temperature versus municipal server counts) yields statistically significant values exceeding despite zero causal or predictive relationship.
Robust time series forecasting requires structuring problems around temporal dependencies: decomposing trends and cycles, enforcing stationarity through differencing, analyzing autocorrelation, and validating strictly using rolling-origin expanding windows.
Think of It Like This
Predicting tomorrow's temperature on an orbiting space station
Imagine you live in a space station rotating once every 24 hours while gradually drifting closer to a star over several months. If you want to forecast tomorrow's temperature, your thermometer reading is not a random draw.
It is the sum of three distinct physical influences:
- A slow, persistent trend (the station creeping closer to the star).
- A predictable, repeating seasonal cycle (the 24-hour day/night thermal rotation).
- Random noise (unpredictable solar flares or air-filter turbulence).
If you want to test whether your forecast algorithm works, you can only give it data from the first two months and ask it to predict the third month. Testing it by giving it random days from month three to predict days in month one is cheating—in real life, tomorrow's sunrise never happens yesterday.
How It Actually Works
Decomposition, Stationarity, and Rolling Horizons
1. Classical Decomposition
A raw univariate time series is modeled as a combination of three structural components:
- Additive Decomposition: Used when the amplitude of seasonal fluctuations remains constant regardless of the series level:
- Multiplicative Decomposition: Used when the seasonal variation scales proportionally with the trend:
The trend-cycle is estimated using a centered moving average of window size equal to the seasonal period . The de-trended series is averaged across repeating periods to isolate , leaving stationary residual noise .
2. Stationarity and Differencing
A time series is weakly stationary (wide-sense stationary) if its statistical properties do not change over time:
- Constant Mean: for all .
- Constant Variance: for all .
- Time-Invariant Autocovariance: , depending solely on the temporal lag , not on the calendar index .
Non-stationary series containing deterministic or stochastic trends are made stationary through differencing:
If a seasonal pattern with period exists (e.g., 7 days for daily data), seasonal differencing removes it:
Stationarity is verified statistically using unit root tests such as the Augmented Dickey-Fuller (ADF) test, which tests the null hypothesis that a unit root is present ( in ). A -value rejects the unit root, confirming stationarity.
3. Autocorrelation Function (ACF) and Partial Autocorrelation (PACF)
The sample autocorrelation at lag measures the linear correlation between observations steps apart:
Values outside the Bartlett confidence bands are statistically significant (). Spikes at lag reveal periodic seasonal cycles and indicate the order and for autoregressive moving-average (ARIMA) models or the optimal lookback window for deep recurrent networks.
4. Leak-Free Rolling-Origin Cross-Validation
Because temporal sequence is unidirectional, standard random cross-validation leaks information. Instead, practitioners use rolling-origin (expanding window) evaluation:
- Fold 1: Train on , test on forecast horizon .
- Fold 2: Train on , test on forecast horizon .
- Fold 3: Train on , test on forecast horizon .
Forecast accuracy across horizons is measured via scale-independent metrics such as Mean Absolute Scaled Error (MASE) or Symmetric Mean Absolute Percentage Error (sMAPE):
A MASE proves the model outperforms a naive persistence baseline (predicting ).
Worked Example
Let us trace autocorrelation and first-order differencing on a 6-day traffic sequence:
-
Calculate Sample Mean :
-
Compute Total Variance Sum of Squared Deviations (Denominator):
-
Compute Lag-1 Covariance (Numerator for ):
-
Calculate Autocorrelation : A high positive value () reflects the persistent upward trend.
-
Apply First Differencing: Differencing eliminates the level shift, converting the raw values into stabilized incremental step rates.
Code
from typing import List, Tuple
def compute_acf(series: List[float], max_lag: int) -> List[float]: """Computes sample autocorrelation function (ACF) up to max_lag.""" n = len(series) mean_val = sum(series) / n denom = sum((x - mean_val) ** 2 for x in series) if denom == 0.0: return [1.0] + [0.0] * max_lag
acf_values = [1.0] # Lag 0 is always 1.0 for k in range(1, max_lag + 1): num = sum((series[t] - mean_val) * (series[t - k] - mean_val) for t in range(k, n)) acf_values.append(num / denom) return acf_values
def difference_series(series: List[float], lag: int = 1) -> List[float]: """Calculates lag-k differencing to remove trend or seasonality.""" return [series[i] - series[i - lag] for i in range(lag, len(series))]
def rolling_window_split( total_len: int, initial_train: int, horizon: int) -> List[Tuple[Tuple[int, int], Tuple[int, int]]]: """Generates strictly leak-free expanding window train/test index bounds.""" splits = [] current_end = initial_train while current_end + horizon <= total_len: train_bounds = (0, current_end) test_bounds = (current_end, current_end + horizon) splits.append((train_bounds, test_bounds)) current_end += horizon return splits
raw_data = [100.0, 105.0, 112.0, 120.0, 129.0, 139.0]acf_vals = compute_acf(raw_data, max_lag=2)diff_data = difference_series(raw_data, lag=1)
print("ACF lags [0, 1, 2]:", [round(v, 4) for v in acf_vals])# -> ACF lags [0, 1, 2]: [1.0, 0.5027, 0.1065]print("Differenced series:", diff_data)# -> Differenced series: [5.0, 7.0, 8.0, 9.0, 10.0]
# Expanding window CV across 10 steps with train=6, horizon=2folds = rolling_window_split(total_len=10, initial_train=6, horizon=2)for idx, (tr, te) in enumerate(folds, 1): print(f"Fold {idx}: Train indices {tr}, Test indices {te}")# -> Fold 1: Train indices (0, 6), Test indices (6, 8)# -> Fold 2: Train indices (0, 8), Test indices (8, 10)Watch Out For
Future target leakage via global feature scaling and standardizers
Symptom: An XGBoost or LSTM forecasting pipeline yields near-zero test error in cross-validation, but error surges by when deployed live.
This silent bug occurs when StandardScaler().fit_transform(X) or MinMaxScaler() is executed over the full dataset before splitting into train and test sets. Fitting scalers globally leaks the future minimum, maximum, mean, and standard deviation into historical training steps. Similarly, creating rolling window features (like a 7-day rolling mean) without setting shift(1) causes the target value at to leak into the feature vector for time .
Always fit all preprocessors, scalers, imputers, and encoders strictly on the training fold slice, then call transform() on the validation/test horizons. Ensure every lagged feature strictly references or earlier: df['lag_1'] = df['y'].shift(1).
The Quick Version
- Time series data breaks the assumption due to strong temporal autocorrelation and chronological ordering.
- Classical decomposition separates series into trend-cycle , repeating seasonal cycle , and residual noise .
- Weak stationarity requires constant mean, constant variance, and autocovariance that depends only on the temporal lag.
- Differencing stabilizes non-stationary means and removes secular polynomial trends.
- Never use random -fold cross-validation; always use expanding rolling-origin splits to prevent temporal leakage.