Skip to content
AI360Xpert
Beta
Time Series & Forecasting

Time Series & Forecasting

30 interview questions in this topic, each with its full answer shown below. Use "Collapse all" to skim just the titles.

“How do you avoid look-ahead bias and leakage when building time series features?”

Quick answer

Ensure all features, especially lags and rolling windows, are calculated strictly using data available prior to the prediction time. Shift features appropriately and use walk-forward validation instead of random cross-validation.

Answer

Look-ahead bias (or data leakage) occurs when information that would not realistically be available at the time of prediction is inadvertently included in the model's training features. In time series forecasting, this is the most common reason for a model performing perfectly in training but failing utterly in production.

How to avoid it:

  1. Strictly Shift Features: When predicting the value at time tt, your features must only be derived from data at t−1,t−2t-1, t-2, etc. If you want a 7-day rolling average to predict tomorrow, you must shift that rolling average by at least one day.
  2. Be Careful with Multi-Step Horizons: If you are forecasting 7 days ahead (t+7t+7) using a direct strategy, you cannot use a lag-1 feature (t+6t+6). At time tt, the most recent data you have is tt. Therefore, to predict t+7t+7, the closest lag you can use is lag-7.
  3. Impute Carefully: Never impute missing values using global statistics. If you replace a missing value on day 10 with the average of the entire 100-day series, you have leaked data from days 11-100 into day 10. Imputation must only use forward-filling or historical windows.
  4. Target Scaling: If you normalize or scale your target variable (e.g., using StandardScaler), you must fit the scaler only on the training split, not on the entire dataset.
  5. No Random Validation: Never use kk-fold cross-validation or random train-test splits. Always use time-based splits or Walk-Forward (Rolling Origin) validation.
  6. Feature calculation limits: When creating aggregated features (like the historical average sales per customer), those aggregates must be computed cumulatively (expanding window) up to time t−1t-1, rather than grouped globally over the whole dataset.

💡 Note: A simple sanity check for data leakage: If you train a model to predict t+1t+1 and it achieves an R2R^2 of 0.99 or a near-zero MAE on complex, noisy business data, you almost certainly have a leakage problem.

“How do you produce calibrated prediction intervals and probabilistic forecasts?”

Quick answer

Use methods like quantile regression to predict specific percentiles (e.g., 10th and 90th) directly, or use conformal prediction to generate statistically guaranteed intervals. Deep learning models like DeepAR output parameters of a probability distribution.

Answer

A point forecast (predicting exactly 100 units of sales) is often insufficient for business decisions like inventory management, which require understanding the uncertainty of the prediction. We need prediction intervals (e.g., "we are 90% confident sales will be between 80 and 120").

Here is how you generate them depending on your modeling approach:

1. Statistical Models (ARIMA, Prophet) Classical models usually output confidence/prediction intervals natively by assuming the residuals (errors) follow a normal distribution. If the residuals are not normally distributed, these intervals will not be mathematically calibrated.

2. Quantile Regression (Tree-based models) Instead of minimizing MSE (which predicts the mean), you can configure models like LightGBM or GradientBoosting to minimize the Pinball Loss function. By doing this, you can train one model to predict the 10th percentile and another to predict the 90th percentile. This produces non-parametric intervals without assuming a normal distribution.

3. Probabilistic Deep Learning (DeepAR) Models like Amazon's DeepAR don't output a single number. Instead, the final layer of the neural network outputs the parameters of a chosen distribution (e.g., the mean μ\mu and standard deviation σ\sigma of a Gaussian, or the parameters of a Negative Binomial). You can then sample from this distribution to get prediction intervals.

4. Conformal Prediction This is a powerful, model-agnostic post-processing technique. It looks at the historical errors your point-forecasting model made on a calibration dataset and uses those empirical errors to draw guaranteed bounds around future predictions, regardless of the underlying model architecture or data distribution.

💡 Note: Do not confuse confidence intervals with prediction intervals. A confidence interval bounds the estimate of a population parameter (like the mean). A prediction interval bounds the value of a single future observation, and must therefore be wider to account for the inherent noise in the data.

“What are common metrics for evaluating forecasts, such as MAE, RMSE, and MAPE?”

Quick answer

Common metrics include MAE (mean absolute error) for robust average error, RMSE (root mean squared error) for penalizing large errors heavily, and MAPE (mean absolute percentage error) for expressing error as a relative percentage of actuals.

Answer

Evaluating the accuracy of a time series forecast requires comparing the predicted values against the actual observed values. Several metrics are commonly used, each with different mathematical properties and use cases:

  1. MAE (Mean Absolute Error):

    • Formula: 1n∑∣yi−y^i∣\frac{1}{n} \sum |y_i - \hat{y}_i|
    • Interpretation: The average absolute magnitude of the errors. It is measured in the same units as the original data.
    • Pros/Cons: Highly interpretable and robust to outliers. However, it doesn't heavily penalize very large, catastrophic errors.
  2. RMSE (Root Mean Squared Error):

    • Formula: 1n∑(yi−y^i)2\sqrt{\frac{1}{n} \sum (y_i - \hat{y}_i)^2}
    • Interpretation: Also measured in the original units, but because the errors are squared before being averaged, it gives disproportionately high weight to large errors.
    • Pros/Cons: Ideal when large errors are particularly costly to your business. It is more sensitive to outliers than MAE.
  3. MAPE (Mean Absolute Percentage Error):

    • Formula: 100%n∑∣yi−y^iyi∣\frac{100\%}{n} \sum |\frac{y_i - \hat{y}_i}{y_i}|
    • Interpretation: Expresses the error as a percentage of the actual value.
    • Pros/Cons: Excellent for communicating with business stakeholders ("our model is off by 5% on average") and for comparing performance across series with different scales. However, MAPE is undefined or infinite if the actual value yiy_i is zero, and it is biased—it penalizes positive errors (over-forecasting) more heavily than negative errors (under-forecasting).

Other useful metrics include sMAPE (symmetric MAPE, which addresses some of MAPE's biases) and MASE (Mean Absolute Scaled Error, which compares the forecast to a naive baseline).

💡 Note: Never rely on a single metric. A model optimized purely for MAE might predict the median, missing all the variance, while a model optimized for RMSE might overreact to outliers.

“How would you detect anomalies in streaming time series with concept drift?”

Quick answer

Use adaptive models that update over time, such as rolling window statistics, exponentially weighted moving averages (EWMA), or online learning algorithms like Half-Space Trees, which automatically adjust their baselines as the underlying data distribution changes.

Answer

Anomaly detection in time series involves identifying data points that deviate significantly from expected behavior. In a streaming environment (where data arrives continuously) with concept drift (where the underlying statistical properties of the data change over time), static models will fail. A model trained on last year's server traffic will flag all of today's traffic as anomalous if your user base has doubled.

To handle concept drift in streaming data, you must use adaptive algorithms that continuously update their internal baselines:

  1. Rolling Window Statistics: The simplest approach. You maintain a moving window of the last NN observations. Calculate the rolling mean and standard deviation. If a new data point is more than 3 standard deviations away from the rolling mean, flag it. As the series drifts naturally, the rolling mean follows it.
  2. EWMA (Exponentially Weighted Moving Average): Similar to the rolling window but applies exponentially decreasing weights to older data. It reacts faster to recent changes and is computationally cheaper because it doesn't require storing a window of data, just the previous state.
  3. Adaptive Forecasting Models: Use a model like Holt-Winters or a streaming ARIMA to predict the next value t+1t+1. Compare the actual value to the prediction. If the residual error is beyond a dynamic threshold, it's an anomaly.
  4. Online Machine Learning: Algorithms like Half-Space Trees (HST) or Loda are designed for streaming environments. They build isolation forests incrementally, constantly updating the tree structures as new data flows in, allowing them to adapt gracefully to concept drift without requiring full retraining.

💡 Note: When an anomaly is detected, it is crucial not to include that massive spike in the baseline update for future steps, otherwise the model's threshold will become distorted and it will miss subsequent anomalies.

“How would you detect change points and structural breaks?”

Quick answer

Use algorithms like PELT or Binary Segmentation to minimize a cost function (like variance) across partitions. You can also use the CUSUM (Cumulative Sum) chart to detect small, persistent shifts in the mean over time.

Answer

A structural break or change point occurs when the fundamental properties of a time series (its mean, variance, or trend) abruptly and persistently change. Unlike an anomaly (a temporary spike), a change point marks a permanent shift in the data generating process—for example, a competitor dropping prices, or a sensor being permanently recalibrated.

If a forecasting model is trained across a structural break without accounting for it, the model will fail.

Methods for detection:

  1. Offline Methods (Historical Data): If you have a fixed dataset and want to find historical breaks, the goal is to partition the data into segments such that the segments are internally consistent but statistically different from each other.
    • PELT (Pruned Exact Linear Time): A highly efficient algorithm that tests different partition points to minimize a cost function (like changes in mean or variance) while applying a penalty for creating too many segments.
    • Binary Segmentation: A greedy algorithm that finds a single change point that minimizes the cost function, then recursively splits the resulting two segments to find further points.
  2. Online Methods (Streaming Data):
    • CUSUM (Cumulative Sum): This is a classic control chart technique. It calculates the cumulative sum of deviations from a target mean. Small random variations cancel each other out, keeping the sum near zero. But a persistent shift in the mean will cause the CUSUM to drift linearly upwards or downwards until it breaches a predefined threshold, signaling a break.
    • Chow Test: An econometrics test used to determine if the coefficients of a linear regression model are different in two different subsets of the data.

💡 Note: Once a structural break is identified in historical data, you typically either truncate the data (only train the model on data after the break) or include a dummy variable to help the model learn the new baseline.

“How would you evaluate whether a forecasting model really beats a naive seasonal baseline?”

Quick answer

Compare the complex model's MAE or RMSE against a seasonal naive baseline (which simply predicts the value from the previous season, like last week). If the complex model's error is not significantly lower, the added complexity isn't justified.

Answer

A common trap in time series forecasting is building a massive, complex machine learning or deep learning model that looks great according to its MAE or MAPE, but actually performs worse than a trivially simple rule.

Before putting any complex model into production, it must prove its worth against naive baselines.

  1. Naive Forecast: The forecast for tomorrow is simply whatever happened today (yt+1=yty_{t+1} = y_t). This is the baseline for random walk data (like stock prices).
  2. Seasonal Naive Forecast: The forecast for tomorrow is simply whatever happened one season ago. If the data has weekly seasonality, the forecast for this Tuesday is whatever happened last Tuesday (yt+1=yt−6y_{t+1} = y_{t-6}).

How to Evaluate: Calculate your error metric (e.g., MAE) for both the complex model and the seasonal naive baseline on your test set. If the complex model's MAE is not significantly lower, the model is not learning anything useful.

A formal way to measure this is the Mean Absolute Scaled Error (MASE). MASE divides the MAE of your model by the MAE of a naive baseline (calculated on the training set).

  • If MASE<1MASE < 1: Your model is doing better than the naive baseline.
  • If MASE>1MASE > 1: Your model is doing worse than the naive baseline, and you should literally just copy-paste last week's data instead of using your model.

💡 Note: In many high-noise business environments (like daily retail sales at a granular store level), a seasonal naive baseline is incredibly hard to beat. Always build this baseline first—it takes 3 lines of code and sets the benchmark for all future modeling efforts.

“Explain ARIMA and what the p, d, and q parameters mean.”

Quick answer

ARIMA is a forecasting model combining Autoregression (AR, parameter 'p' for lags), Integration (I, parameter 'd' for differencing to achieve stationarity), and Moving Average (MA, parameter 'q' for lagged forecast errors).

Answer

ARIMA stands for AutoRegressive Integrated Moving Average. It is a widely used statistical method for analyzing and forecasting time series data. The model is defined by three distinct components, governed by three non-negative integer parameters: (p,d,q)(p, d, q).

  1. AR (AutoRegressive) - pp: The autoregressive part indicates that the evolving variable of interest is regressed on its own lagged (i.e., prior) values. The parameter pp specifies the number of lag observations included in the model. If p=2p=2, the model uses the values at time t−1t-1 and t−2t-2 to predict time tt.

  2. I (Integrated) - dd: The integrated part represents the differencing of raw observations to allow the time series to become stationary (i.e., data values are replaced by the difference between the data values and the previous values). The parameter dd represents the degree of differencing. If d=1d=1, you are modeling the first difference of the series (yt−yt−1y_t - y_{t-1}). If the series is already stationary, dd is 0.

  3. MA (Moving Average) - qq: The moving average part indicates that the regression error is actually a linear combination of error terms whose values occurred contemporaneously and at various times in the past. The parameter qq specifies the size of the moving average window (order of moving average). It models the next value based on the residual errors from previous predictions.

In mathematical terms, an ARIMA(1,1,1) model says that the change in the series today is a function of the change yesterday, plus the error made yesterday, plus some baseline noise.

💡 Note: Finding the correct values for p,d,qp, d, q often involves examining ACF and PACF plots, or using automated grid search algorithms like auto.arima which iterate through combinations and select the model with the lowest AIC or BIC score.

“Explain state-space models and the Kalman filter.”

Quick answer

State-space models represent a time series as an unobserved 'hidden state' that evolves over time, producing the observations we see. The Kalman filter is a recursive algorithm that optimally estimates this hidden state from noisy observations.

Answer

State-Space Models (SSMs) offer a highly flexible, probabilistic framework for modeling time series data. In this framework, the time series you actually observe (YtY_t) is assumed to be an imperfect, noisy measurement of an underlying, unobserved "hidden state" (XtX_t) that is evolving over time.

An SSM is defined by two equations:

  1. The State Equation (Transition Equation): Describes how the hidden state evolves from time t−1t-1 to tt. (e.g., A car's actual position today is its position yesterday plus its velocity, plus some random wind/friction noise).
  2. The Observation Equation (Measurement Equation): Describes how the hidden state translates into the data we actually observe. (e.g., The GPS reading we get today is the car's actual position plus satellite measurement error).

The Kalman Filter is the mathematical algorithm used to estimate the true hidden state XtX_t given the noisy observations Y1…YtY_1 \dots Y_t. It works recursively in two steps:

  1. Predict: It uses the state equation and the previous state estimate to predict the current state, before seeing the new observation.
  2. Update: It takes the new observation and updates the prediction. The Kalman filter calculates a "Kalman Gain," which optimally weighs how much to trust the model's prediction versus how much to trust the new observation (based on their respective variances).

Virtually all classical time series models (including ARIMA and Holt-Winters) can be rewritten as special cases of state-space models. They are particularly powerful for structural time series modeling (like Prophet), handling missing data (the filter just runs the 'Predict' step without the 'Update' step), and combining multiple sensor inputs.

💡 Note: The standard Kalman filter assumes linear transitions and Gaussian noise. If the system is non-linear, you must use an Extended Kalman Filter (EKF) or a Particle Filter.

Quick answer

Instead of training a separate model per series, use a global forecasting model (like LightGBM or DeepAR) trained on panel data. This allows the model to learn cross-series patterns, handle cold starts, and utilize shared features like holidays or store attributes.

Answer

When dealing with high-dimensional forecasting problems—such as predicting sales for 10,000 different SKUs across 500 different stores—the classical approach of training a single ARIMA or Prophet model for each individual series (the local model approach) scales poorly and misses valuable information.

The modern best practice is to build a global model.

In a global model approach, you stack all your individual time series vertically into a single large tabular dataset (often called panel data). You then train a powerful machine learning model (like LightGBM, XGBoost, or a deep learning model like Amazon's DeepAR) on this combined dataset.

Advantages of Global Models:

  1. Cross-Learning: The model learns shared patterns across different series. If it learns how ice cream sales behave during a heatwave in Store A, it can apply that knowledge to Store B, even if Store B has less historical data.
  2. Cold Starts: If you introduce a brand new SKU with no history, a local ARIMA model cannot forecast it. A global model can forecast it immediately by relying on static features (e.g., item category, price, brand) and looking at how similar items behaved in the past.
  3. Fewer Models to Maintain: Managing one large LightGBM model in production is far easier than orchestrating the training and deployment of 5,000,000 individual statistical models.
  4. Feature Engineering: You can easily include global features (like national holidays or macroeconomic indicators) and static features (like store square footage).

💡 Note: To make a global model work, you must include categorical identifiers (like store_id and sku_id) as features. You also must be extremely careful to normalize or scale the target variable if different series have vastly different magnitudes.

“How do you handle intermittent demand with many zeros?”

Quick answer

Use models specifically designed for sparse data, such as Croston's method (which separates demand timing and demand size), zero-inflated Poisson/Negative Binomial regression, or aggregate the data to a lower frequency (e.g., weekly instead of daily).

Answer

Intermittent demand occurs when a time series contains a large number of zero values, punctuated by sporadic, irregular non-zero values (e.g., daily sales of spare aviation parts or high-end luxury goods).

Standard forecasting models (like ARIMA or simple exponential smoothing) fail completely on this data. They will usually predict a continuous, fractional, low-level constant (e.g., predicting 0.15 sales every day), which is practically useless for inventory planning.

Here is how to handle it:

1. Croston's Method This is the classical approach for intermittent demand. Instead of forecasting the series directly, Croston's method splits the problem into two separate exponential smoothing models:

  • One model forecasts the time interval between non-zero demands.
  • One model forecasts the size of the demand, given that a demand occurs. The final forecast is the ratio of the demand size to the interval. (Variations like Syntetos-Boylan Approximation - SBA - correct for mathematical biases in Croston's original formula).

2. Zero-Inflated Count Models Because demand is usually discrete (you can't sell 0.5 parts), use statistical models designed for count data, such as Zero-Inflated Poisson or Zero-Inflated Negative Binomial regression. These models assume the data is generated by two processes: one determining if the count is zero, and another determining the count size if it's non-zero.

3. Temporal Aggregation Often, the simplest solution is to change the frequency of the data. If a series is heavily intermittent at a daily level, aggregating it to a weekly or monthly level will often eliminate the zeros, resulting in a smooth, continuous series that standard models can handle perfectly.

💡 Note: For intermittent demand, standard error metrics like MAPE are useless because you will be dividing by zero. Use metrics like MASE (Mean Absolute Scaled Error) or wMAPE (weighted MAPE) instead.

“How do you handle missing values and irregular timestamps?”

Quick answer

Irregular timestamps must be resampled to a fixed frequency. Missing values can be handled via forward-filling, interpolation, or seasonal imputation, depending on the gap size and data characteristics.

Answer

Time series models, especially classical statistical models and models relying on lag features, strictly require a continuous, equally spaced sequence of observations.

Handling Irregular Timestamps: If data arrives at random intervals (e.g., website clicks or sensor events), it must be resampled to a fixed frequency. In pandas, you would use .resample('H') for hourly or .resample('D') for daily data.

  • Downsampling (high frequency to low frequency): Requires an aggregation function, such as sum() for sales volume or mean() for temperature.
  • Upsampling (low frequency to high frequency): Will introduce missing values that must be imputed.

Handling Missing Values (Imputation): Once the frequency is fixed, you must deal with the resulting NaN values. The choice of imputation depends entirely on the nature of the data:

  1. Forward Fill (Last Observation Carried Forward - LOCF): Propagates the last known value forward. Ideal for state-based metrics (e.g., the price of a stock remains the same until a new trade happens).
  2. Linear Interpolation: Draws a straight line between the points before and after the missing gap. Good for continuous physical measurements (e.g., temperature) when the gap is very short.
  3. Spline/Polynomial Interpolation: Fits a smooth curve. Can be dangerous if the gap is large, as polynomials can swing wildly.
  4. Seasonal Imputation: Replaces the missing value with the historical average of that specific season (e.g., replacing a missing December sales figure with the average of previous Decembers).
  5. Model-based Imputation: Using a separate machine learning model, or a tool like KNN, to predict the missing values based on other correlated series.

💡 Note: Never use the global mean of the entire time series to impute missing values, as this destroys local trends and seasonal patterns, and causes data leakage if the mean includes future data.

“How would you do hierarchical forecasting and reconcile forecasts across levels?”

Quick answer

Hierarchical forecasting predicts data structured in levels (e.g., country > state > store). Reconciliation ensures the forecasts add up correctly. Bottom-up sums lower-level forecasts; Top-down distributes the top forecast; Optimal reconciliation adjusts all levels simultaneously based on variance.

Answer

Many business time series are structurally hierarchical. For example, product sales can be aggregated from SKU level →\rightarrow Brand level →\rightarrow Category level →\rightarrow Total Company level.

If you forecast the SKU level separately from the Category level, the sum of your SKU forecasts will almost certainly not match your Category forecast. Reconciliation is the mathematical process of adjusting the base forecasts so they perfectly aggregate.

There are several methods for hierarchical reconciliation:

  1. Bottom-Up: You only generate forecasts for the most granular level (the SKUs). You then sum them up to get the higher-level forecasts.
    • Pros: Ensures perfect reconciliation; captures unique dynamics of individual series.
    • Cons: Granular series are often very noisy and sparse, leading to poor individual forecasts that accumulate errors when summed.
  2. Top-Down: You only forecast the aggregate top level (Total Company). You then disaggregate this down to the lower levels based on historical proportions (e.g., SKU AA usually accounts for 5% of Total Sales).
    • Pros: Top-level data is usually very smooth and easy to forecast accurately.
    • Cons: Completely misses changing dynamics at the bottom level (e.g., SKU AA is growing, SKU BB is dying).
  3. Middle-Out: A hybrid where you forecast an intermediate level (Brand), sum up for higher levels, and distribute down for lower levels.
  4. Optimal Reconciliation (MinT - Minimum Trace): You forecast every level of the hierarchy independently. You then use linear algebra and the covariance matrix of the forecast errors to optimally adjust all the forecasts. It intelligently shifts the forecasts, trusting the levels that have lower historical variance more than the noisy levels.

💡 Note: Optimal reconciliation (like the MinT approach) is generally considered the state-of-the-art because it leverages information from all levels simultaneously, consistently outperforming top-down and bottom-up approaches.

“How do ACF and PACF plots help choose model orders?”

Quick answer

ACF (Autocorrelation Function) plots help determine the MA (q) order by showing where correlations cut off, while PACF (Partial Autocorrelation Function) plots help determine the AR (p) order by isolating direct correlations at specific lags.

Answer

Autocorrelation Function (ACF) and Partial Autocorrelation Function (PACF) plots are standard visual tools used to determine the order of AutoRegressive (AR, pp) and Moving Average (MA, qq) terms in an ARIMA model.

Before using these plots, the time series must be stationary. If it is not, apply differencing (dd) until it is.

1. ACF (Autocorrelation Function): The ACF plot shows the correlation between a series and its lags. It captures both the direct effect of lag t−kt-k on tt, and the indirect effect propagated through the intervening lags (e.g., t−kt-k affects t−1t-1, which affects tt).

  • Use: Identifying the MA (qq) order.
  • Rule: If the ACF plot cuts off sharply after lag qq, and the PACF plot tails off (decays gradually), the series likely requires an MA(qq) model. The sharp cut-off indicates that the error term dependency only lasts for qq periods.

2. PACF (Partial Autocorrelation Function): The PACF plot shows the correlation between a series and its lags after removing the effects of intervening lags. It isolates the pure, direct relationship between yty_t and yt−ky_{t-k}.

  • Use: Identifying the AR (pp) order.
  • Rule: If the PACF plot cuts off sharply after lag pp, and the ACF plot tails off, the series likely requires an AR(pp) model. The sharp cut-off in PACF indicates that lag pp is the furthest back in time that has a direct predictive effect.

If both ACF and PACF tail off gradually, the model likely requires both AR and MA terms (a mixed ARIMA model). If neither tails off, the series may not be stationary yet.

💡 Note: In modern practice, while ACF/PACF visual inspection is useful for intuition and understanding the data, most practitioners rely on hyperparameter search algorithms (like pmdarima in Python) to find the optimal (p,d,q)(p,d,q) parameters by minimizing the Akaike Information Criterion (AIC).

“How do you test for stationarity, and how do you make a series stationary?”

Quick answer

Test for stationarity using visual inspection, summary statistics over partitions, and statistical tests like the Augmented Dickey-Fuller (ADF) test. Make a series stationary by differencing, taking logarithms, or using power transformations.

Answer

Determining if a time series is stationary is a crucial first step for many classical forecasting models.

How to test for stationarity:

  1. Visual Inspection: Plot the time series. Look for clear upward or downward trends, changing variance (the peaks and troughs getting wider over time), or obvious seasonality. If you see these, the series is not stationary.
  2. Summary Statistics: Split your data into two or three contiguous chunks. Calculate the mean and variance for each chunk. If these values are significantly different across the chunks, the series is likely non-stationary.
  3. Statistical Tests: The most common rigorous method is the Augmented Dickey-Fuller (ADF) test. The null hypothesis of the ADF test is that the time series has a unit root (meaning it is non-stationary). If the p-value is less than your significance level (e.g., <0.05< 0.05), you reject the null hypothesis and conclude the series is stationary. Another popular test is the KPSS test, which has the opposite null hypothesis.

How to make a series stationary:

  1. Differencing: To remove a trend, subtract the previous observation from the current observation (yt′=yt−yt−1y_t' = y_t - y_{t-1}). If the trend is non-linear, you may need a second difference. To remove seasonality, use seasonal differencing (e.g., for monthly data, subtract the value from 12 months ago: yt′=yt−yt−12y_t' = y_t - y_{t-12}).
  2. Log Transform: If the variance of the series increases with time (heteroscedasticity), applying a natural logarithm can stabilize the variance.
  3. Power Transforms: Box-Cox or Yeo-Johnson transformations can also stabilize variance and make the data more normally distributed.

💡 Note: You will often need to combine these techniques. For example, a series with an exponentially growing trend and increasing seasonal variance will usually require a log transform followed by differencing to become stationary.

“How does walk-forward (rolling origin) validation work?”

Quick answer

Walk-forward validation evaluates time series models by repeatedly training on historical data and predicting the immediate future. After each step, the origin shifts forward, incorporating new actuals into the training set for the next prediction.

Answer

Because random cross-validation destroys the chronological order of time series data and causes data leakage, forecasting models must be evaluated using walk-forward validation (also known as rolling origin validation, rolling forecasting origin, or time series backtesting).

Walk-forward validation simulates how a forecasting model is actually used in production: predicting the future, waiting to see the actual results, retraining or updating the model with those new actuals, and predicting again.

Here is how the process works:

  1. Initial Split: Define a minimum historical window required to train the model. For example, if you have 100 days of data, use days 1-80 as the initial training set.
  2. Forecast: Train the model on days 1-80 and forecast a specific horizon, say, the next 3 days (days 81-83).
  3. Evaluate: Compare the forecast to the actual values for days 81-83 and record the error metric.
  4. Step Forward (Rolling the Origin): Shift the training window forward. Depending on your strategy, you can either:
    • Expand the window: Train on days 1-83.
    • Slide the window: Train on days 4-83 (keeping the window size fixed at 80).
  5. Repeat: Forecast the next 3 days (days 84-86), evaluate, and step forward again. Repeat this process until you reach the end of your dataset.

Finally, aggregate the error metrics across all the validation folds to get a robust estimate of the model's true performance.

💡 Note: The step size (how many periods you move forward each time) and the forecast horizon (how far out you predict) should match your actual business deployment cycle. If you deploy weekly to forecast the next 14 days, your walk-forward validation should step by 7 days and forecast 14 days.

“How would you incorporate holidays and external regressors?”

Quick answer

Holidays are typically incorporated as binary indicator variables or dummy variables covering the days before and after the event. External regressors are added as additional features (in ML models) or as exogenous variables (in ARIMAX/SARIMAX).

Answer

Real-world time series, particularly in retail and economics, are heavily influenced by external factors that cannot be captured purely by looking at past values.

Incorporating Holidays: Holidays often create massive, predictable spikes or dips in data. Because they don't always fall on the exact same date (e.g., Easter moves, Thanksgiving is the 4th Thursday), simple seasonality models fail to capture them.

  • In machine learning models (like XGBoost), holidays are typically added as binary indicator columns (e.g., is_christmas: 1/0).
  • Because holiday effects often bleed into surrounding days, it is highly recommended to engineer features like days_until_christmas or create window features (e.g., a dummy variable covering the week leading up to the holiday).
  • Models like Prophet handle holidays natively by requiring a dataframe of dates and holiday names, and they automatically learn a parameter for the "bump" associated with each holiday.

Incorporating External Regressors: External regressors (or exogenous variables) are independent variables that affect your target variable. For example, if you are forecasting ice cream sales, temperature is a powerful external regressor.

  • In statistical models, you upgrade ARIMA/SARIMA to ARIMAX/SARIMAX. The "X" stands for eXogenous. The model basically performs a linear regression on the external variables and uses the ARIMA portion to model the residuals (the errors).
  • In machine learning models, external regressors are simply added as additional columns in your feature set.

💡 Note: A critical requirement for using external regressors in forecasting is that you must know the future value of the regressor at the time you make the forecast. For holidays, this is easy (the calendar is known). For weather, you must rely on weather forecasts, which introduces secondary forecast error into your primary model.

“How do LSTM, Temporal Convolutional Networks, and Transformers compare for time series forecasting?”

Quick answer

LSTMs capture sequential dependencies but struggle with very long contexts. TCNs use dilated convolutions to capture long sequences efficiently. Transformers use attention to model global dependencies but can be data-hungry and prone to overfitting on small datasets.

Answer

Deep learning has become increasingly popular for time series forecasting, especially for massive multivariate datasets. The three dominant architectures each handle temporal dynamics differently:

1. LSTMs (Long Short-Term Memory)

  • Mechanism: A type of Recurrent Neural Network (RNN) that processes data sequentially step-by-step, maintaining a hidden "state" that carries information across time.
  • Pros: Naturally suited for sequential data. Excellent at capturing short-to-medium term temporal dependencies.
  • Cons: Slow to train because they cannot be parallelized (step tt requires step t−1t-1 to finish). They suffer from the vanishing gradient problem over very long sequences, often "forgetting" distant past information.

2. TCNs (Temporal Convolutional Networks)

  • Mechanism: Adapts CNNs for time series using 1D causal convolutions (ensuring no future leakage) and dilated convolutions (skipping steps to exponentially increase the receptive field).
  • Pros: Highly parallelizable (fast training). Can maintain a massive receptive field to look far back in time without losing historical context. Often outperforms LSTMs in both speed and accuracy.
  • Cons: Can struggle to model strict, rigid periodicities (like exact daily seasonality) compared to dedicated seasonal models.

3. Transformers (e.g., Informer, Autoformer)

  • Mechanism: Relies on self-attention mechanisms to weigh the importance of all past time steps simultaneously, regardless of their distance from the current step.
  • Pros: Exceptional at capturing long-range global dependencies. Highly parallelizable. They are currently state-of-the-art for many complex, long-horizon multi-step forecasting tasks.
  • Cons: Extremely data-hungry. Prone to severe overfitting if the dataset is small or noisy. The standard attention mechanism has O(N2)O(N^2) complexity relative to sequence length, which requires specialized architectures (like Informer) to compute efficiently.

💡 Note: Despite the hype around deep learning, for univariate forecasting or datasets with less than a few thousand rows, classical methods (ARIMA) or tree-based models (XGBoost) will often outperform these deep architectures while requiring a fraction of the compute.

“What is the difference between one-step and multi-step forecasting? Compare recursive and direct strategies.”

Quick answer

One-step forecasting predicts the immediate next value. Multi-step predicts further into the future. Recursive strategies feed one-step predictions back as inputs for the next step, while direct strategies build a separate model for each future horizon step.

Answer

One-step forecasting predicts only the very next time step (t+1t+1) using all available historical data up to time tt.

Multi-step forecasting predicts a sequence of future values (t+1,t+2,…,t+ht+1, t+2, \dots, t+h), where hh is the forecast horizon. When using models like gradient boosting that expect tabular data with lag features, multi-step forecasting presents a structural challenge: to predict day t+3t+3 using a lag-1 feature, you need the value of day t+2t+2, which hasn't happened yet.

There are two primary strategies to solve this:

1. Recursive Strategy: You build a single, one-step model. To forecast t+1t+1, you use actual historical data. To forecast t+2t+2, you take the prediction you just made for t+1t+1 and plug it back into the model as if it were actual data. You repeat this loop until you reach horizon hh.

  • Pros: Simple to implement; requires only one model.
  • Cons: Error accumulation. If your prediction for t+1t+1 is slightly off, that error becomes the input for t+2t+2, compounding the error as you forecast further into the future.

2. Direct Strategy: You build a separate, distinct model for each step of the horizon. If you want to forecast 7 days out, you train 7 different models. Model 1 is trained to predict t+1t+1 using data up to tt. Model 7 is trained to predict t+7t+7 using data up to tt.

  • Pros: No error accumulation. Model 7 can learn features specifically optimized for a 7-day horizon.
  • Cons: Computationally expensive (you have to train and maintain hh models). It also ignores the inherent correlation between the steps (e.g., t+2t+2 is usually related to t+1t+1).

💡 Note: Modern deep learning architectures (like Sequence-to-Sequence models or Transformers) use a MIMO (Multi-Input Multi-Output) strategy, predicting the entire sequence t+1…t+ht+1 \dots t+h simultaneously in a single forward pass, neatly avoiding the downsides of both recursive and direct methods.

“How do Prophet, ARIMA, and gradient boosting compare for forecasting?”

Quick answer

ARIMA is a classical statistical model requiring stationarity. Prophet is a robust additive regression model built for business time series with holidays. Gradient boosting (e.g., XGBoost) requires heavy feature engineering but often achieves the highest accuracy, especially with many exogenous variables.

Answer

Choosing the right forecasting model depends heavily on the nature of the data and the business context. Here is how three major approaches compare:

1. ARIMA (and SARIMA)

  • How it works: A classical statistical approach that models the temporal dependence (autocorrelation) in stationary data.
  • Pros: Highly interpretable, rigorously grounded in statistical theory, and excellent for short-term forecasting on datasets with strong internal structure.
  • Cons: Requires strict preprocessing (differencing to achieve stationarity). Struggles to incorporate many external regressors efficiently. Fails on non-linear trends.

2. Facebook Prophet

  • How it works: An additive regression model where non-linear trends are fit with yearly, weekly, and daily seasonality, plus holiday effects.
  • Pros: Handles missing data and outliers incredibly well. Requires almost no preprocessing or feature engineering. Excels at "business" time series (e.g., daily sales with strong weekly and yearly seasonality and holiday bumps). Easy to tune for non-experts.
  • Cons: It is fundamentally a curve-fitting tool, not an autoregressive one. It often misses subtle temporal dynamics (like momentum) that ARIMA would catch. Often underperforms more complex models in massive-scale forecasting competitions (like M5).

3. Gradient Boosting (XGBoost, LightGBM, CatBoost)

  • How it works: A tree-based machine learning approach. Time series data must be converted into tabular format using extensive feature engineering (lags, rolling means, date parts).
  • Pros: Dominates forecasting competitions like Kaggle and M5. Easily handles hundreds of exogenous variables, non-linear relationships, and interactions between features. Can forecast thousands of series simultaneously using a single global model.
  • Cons: Cannot extrapolate trends (tree models cannot predict values outside their training domain). Requires massive amounts of manual feature engineering and careful cross-validation to avoid data leakage.

💡 Note: For a single series with strong autoregressive properties, use ARIMA. For a messy business metric with holidays, use Prophet. For large-scale data with many exogenous features, use Gradient Boosting.

“What are trend, seasonality, and noise?”

Quick answer

Trend is the long-term progression of the series, seasonality represents repeating patterns at fixed intervals, and noise is the random, irregular variation left after extracting trend and seasonality.

Answer

In time series analysis, a dataset is often conceptually decomposed into three primary components: trend, seasonality, and noise (or residuals). Understanding these components is critical for building accurate forecasting models.

  1. Trend: The trend is the long-term progression or underlying direction of the data over time. It represents a persistent, systemic upward or downward movement. For instance, global temperatures over decades show an upward trend, while the cost of computing power shows a downward trend. Trends do not have to be linear; they can be exponential, logarithmic, or polynomial.
  2. Seasonality: Seasonality refers to predictable, repeating patterns or cycles that occur at fixed, known intervals. These variations are often tied to the calendar or the clock. Examples include increased retail sales every December (annual seasonality) or peaks in website traffic at 8 PM every day (daily seasonality). Seasonality is distinct from cyclic behavior, which represents rises and falls that are not of a fixed period (like business cycles).
  3. Noise (or Irregularity): Noise is the random, unpredictable variation in the data that remains after the trend and seasonal components have been removed. It is the residual error caused by unpredictable, extraneous factors. In an ideal additive time series decomposition, Value = Trend + Seasonality + Noise. If the noise component contains predictable patterns, the model has failed to capture some signal.

By isolating these components, practitioners can better understand historical behaviors and model them individually before recombining them to produce a forecast.

💡 Note: Decomposition can be additive (Yt=Tt+St+RtY_t = T_t + S_t + R_t) when the seasonal variation is relatively constant, or multiplicative (Yt=Tt×St×RtY_t = T_t \times S_t \times R_t) when the seasonal variation increases with the level of the trend.

“What is the difference between univariate and multivariate time series?”

Quick answer

A univariate time series involves a single variable changing over time, while a multivariate time series involves two or more variables that change together, allowing models to capture complex interactions between them.

Answer

A univariate time series consists of a single sequence of observations over time. You are tracking just one variable. For example, if you are recording the daily closing price of a single stock, or the total daily sales of a store, you are dealing with univariate data. In this scenario, models (like standard ARIMA) rely solely on the historical values of that single variable to forecast its future.

A multivariate time series, on the other hand, consists of two or more variables that are observed simultaneously over time. The key aspect of multivariate series is that the variables typically influence each other. For example, forecasting ice cream sales might involve historical sales data, but also historical daily temperature, humidity, and whether it was a weekend.

In a multivariate setting, models (like VAR - Vector AutoRegression, or modern deep learning architectures) attempt to capture not only the temporal dependencies of a variable on its own past but also the cross-dependencies on the past values of the other variables in the system.

Multivariate forecasting can be much more accurate because it leverages additional context (exogenous variables or co-evolving targets). However, it is also much more complex to model, requires more data, and can suffer from the curse of dimensionality.

💡 Note: When performing multivariate forecasting where you predict variable YY using variable XX, you must often forecast the future values of XX as well, unless XX is deterministic (like day of the week).

“What is a lag feature?”

Quick answer

A lag feature is a variable that contains the value of a time series at a previous time step. It is used to transform time series data into a tabular format, allowing machine learning models to learn temporal dependencies.

Answer

A lag feature is a fundamental tool in time series feature engineering. It involves shifting a time series backward in time so that a previous value becomes a predictor for the current value.

For instance, if your target variable YtY_t is today's sales, a "lag 1" feature would be yesterday's sales (Yt−1Y_{t-1}), and a "lag 7" feature would be sales from exactly one week ago (Yt−7Y_{t-7}). By creating these lag features, you effectively convert a sequential time series problem into a standard supervised learning format (tabular data). This allows non-temporal machine learning models, like Random Forests or Gradient Boosting Machines (XGBoost, LightGBM), to learn the autocorrelation present in the data.

Choosing the right lags is critical. If your data exhibits weekly seasonality, a lag-7 feature is often highly predictive. Autocorrelation (ACF) and Partial Autocorrelation (PACF) plots are standard tools for identifying which lags have the strongest statistical relationship with the current value.

When creating lag features, the first few rows of your dataset will inevitably contain missing values (NaNs) because there is no historical data available prior to the start of the series. These rows must be dropped or imputed before training.

💡 Note: In a multi-step forecasting scenario where you need to predict multiple days into the future at once, you can only use lag features that will actually be known at prediction time. Using a lag-1 feature to predict day 7 requires recursive forecasting.

“What is a moving average?”

Quick answer

A moving average is a technique to smooth data by calculating the average of a fixed window of consecutive points. It helps to highlight underlying trends by filtering out short-term noise.

Answer

A moving average (often called a rolling mean) is a widely used statistical technique for analyzing time series data. It works by creating a series of averages of different subsets of the full data set.

Given a series of numbers and a fixed subset size (often called the "window size" or "lookback period"), the first element of the moving average is obtained by taking the average of the initial fixed subset. Then, the subset is modified by "shifting forward"—excluding the first number of the previous subset and including the next number in the original series. This process is repeated over the entire time series.

The primary purpose of a moving average is to smooth out short-term fluctuations and highlight longer-term trends or cycles. It effectively acts as a low-pass filter, dampening the "noise" in the data.

Common variations include:

  • Simple Moving Average (SMA): The unweighted mean of the previous nn data points.
  • Weighted Moving Average (WMA): Assigns different weights to the data points, usually giving more importance to recent observations.
  • Exponential Moving Average (EMA): Applies exponentially decreasing weights to older observations, reacting more quickly to recent price changes than the SMA.

Moving averages are not only used for data visualization and smoothing but are also foundational to many forecasting algorithms and trading strategies.

💡 Note: While a moving average smooths data, it inherently introduces a "lag" because it is based on past data. A larger window size results in a smoother curve but a greater lag in responding to recent structural changes.

“What is a time series, and how does it differ from regular tabular data?”

Quick answer

A time series is a sequence of data points indexed in time order. Unlike regular tabular data where observations are assumed to be independent, time series data has an inherent temporal order and dependence between successive observations.

Answer

A time series is a sequence of data points recorded or measured at successive, typically evenly spaced, points in time. Examples include daily stock prices, hourly temperature readings, and monthly sales figures.

The fundamental difference between time series data and regular tabular (cross-sectional) data lies in the assumption of independence. In traditional machine learning with tabular data, we often assume that observations are independent and identically distributed (i.i.d.). This means the order of rows does not matter, and shuffling the dataset has no impact on the underlying patterns.

In time series data, the temporal order is crucial. An observation at time tt is typically dependent on observations at time t−1t-1, t−2t-2, and so on. This serial dependence (or autocorrelation) is the very signal we try to capture in forecasting models. If you shuffle the rows of a time series, you destroy the relationship between past and future events, rendering the data useless for forecasting.

Furthermore, time series data is often characterized by specific temporal structures, such as trends (long-term direction) and seasonality (repeating patterns over fixed periods). Because of this temporal dependence, standard validation techniques like random k-fold cross-validation are inappropriate, and time-aware splitting methods like walk-forward validation must be used instead.

💡 Note: When converting time series data to a tabular format for machine learning (e.g., gradient boosting), you must explicitly engineer the temporal dependence into the features using lags, rolling window statistics, and time-based indicators.

“What is autocorrelation?”

Quick answer

Autocorrelation measures the linear relationship between a time series and a lagged version of itself over successive time intervals, helping to identify repeating patterns like seasonality.

Answer

Autocorrelation, also known as serial correlation, is a mathematical representation of the degree of similarity between a given time series and a lagged version of itself over successive time intervals. It measures how strongly the current value of a series depends on its past values.

Just as regular correlation measures the linear relationship between two different variables (e.g., XX and YY), autocorrelation measures the relationship between a variable at time tt (YtY_t) and the same variable at a previous time t−kt-k (Yt−kY_{t-k}), where kk is the lag.

The autocorrelation function (ACF) computes this correlation for various lags kk. The values range from -1 to 1:

  • 1 indicates perfect positive correlation (a high value is followed by a high value).
  • -1 indicates perfect negative correlation (a high value is followed by a low value).
  • 0 indicates no linear relationship.

Autocorrelation is essential for diagnosing time series data:

  1. Identifying Seasonality: If a dataset has monthly seasonality, you will see a strong positive autocorrelation at lag 12, indicating that the value in January of this year is highly correlated with January of last year.
  2. Model Selection: Autocorrelation plots (ACF) and partial autocorrelation plots (PACF) are critical tools for determining the appropriate parameters for autoregressive (AR) and moving average (MA) models.

💡 Note: Autocorrelation only captures linear dependencies. It is possible for two points in a time series to have a strong non-linear relationship even if their autocorrelation is near zero.

“What is exponential smoothing?”

Quick answer

Exponential smoothing is a forecasting method that applies exponentially decreasing weights to older observations. It emphasizes recent data more than older data, making it highly adaptive to recent changes in a time series.

Answer

Exponential smoothing is a family of univariate time series forecasting methods that produce forecasts by calculating weighted averages of past observations. Unlike a simple moving average where all past observations in a fixed window are weighted equally, exponential smoothing assigns exponentially decreasing weights as the observations get older.

The fundamental idea is that recent history is usually more relevant for forecasting the immediate future than the distant past.

In Simple Exponential Smoothing (SES)—used for data with no clear trend or seasonality—the forecast for the next period, y^t+1\hat{y}_{t+1}, is calculated as:

y^t+1=αyt+(1−α)y^t\hat{y}_{t+1} = \alpha y_t + (1 - \alpha) \hat{y}_t

Where:

  • yty_t is the actual observation at time tt.
  • y^t\hat{y}_t is the forecast made for time tt.
  • α\alpha (alpha) is the smoothing parameter, a value between 0 and 1.

If α\alpha is close to 1, the model places almost all weight on the most recent observation, making the forecast highly reactive. If α\alpha is close to 0, the model places weight on a long history of observations, resulting in a much smoother, slower-to-react forecast.

Exponential smoothing is computationally efficient, easy to update as new data arrives, and forms the basis for more advanced models that can handle trend (Holt's Linear Trend method) and seasonality (Holt-Winters method).

💡 Note: The optimal value for the smoothing parameter α\alpha is typically found by minimizing a loss function, such as the Sum of Squared Errors (SSE), over the training data.

“What is the Holt-Winters method?”

Quick answer

The Holt-Winters method is an extension of exponential smoothing that captures level, trend, and seasonality using three distinct smoothing equations. It can handle both additive and multiplicative seasonality.

Answer

The Holt-Winters method, also known as Triple Exponential Smoothing, is a widely used forecasting algorithm that builds upon simple exponential smoothing. While simple exponential smoothing only models the overall level of the data, Holt-Winters breaks the data down into three distinct components, smoothing each one separately.

It uses three smoothing equations (and three corresponding smoothing parameters, α,β,γ\alpha, \beta, \gamma, constrained between 0 and 1):

  1. Level Equation (α\alpha): Smooths the overall baseline value of the series, adjusting for the trend and seasonal components.
  2. Trend Equation (β\beta): Smooths the additive trend (the rate of change), estimating how much the level changes from one step to the next.
  3. Seasonal Equation (γ\gamma): Smooths the seasonal component, isolating the repeating pattern over a specified period (e.g., s=12s=12 for monthly data).

The final forecast is produced by recombining these three components.

A major advantage of the Holt-Winters method is that it can be applied in two variations:

  • Additive Method: Used when the seasonal variations are roughly constant throughout the series. The seasonal component is added to the level and trend.
  • Multiplicative Method: Used when the seasonal variations change proportionally to the level of the series (e.g., as total sales grow over the years, the December holiday spike gets proportionally larger). The seasonal component is multiplied by the level and trend.

💡 Note: Because Holt-Winters requires estimating initial states for the level, trend, and all seasonal indices, it requires a minimum of two full seasonal cycles of historical data to fit accurately.

“What is SARIMA, and when is it needed?”

Quick answer

SARIMA (Seasonal ARIMA) extends ARIMA by adding seasonal terms (P, D, Q, s). It is needed when a time series exhibits repeating seasonal patterns, allowing the model to difference and capture autoregressive relationships across seasonal lags.

Answer

SARIMA stands for Seasonal AutoRegressive Integrated Moving Average. It is an extension of the standard ARIMA model designed specifically to handle time series data that exhibits clear seasonality—repeating patterns at fixed intervals, such as monthly, quarterly, or daily.

Standard ARIMA models (p,d,q)(p, d, q) only look at consecutive data points (e.g., modeling today based on yesterday and the day before). If you have daily retail sales data, sales on a Monday are often strongly correlated with sales on the previous Monday, rather than Sunday. Standard ARIMA struggles to capture this effectively without using an unwieldy number of lag parameters.

SARIMA solves this by adding a second set of parameters to model the seasonal component: SARIMA(p,d,q)(P,D,Q,s)SARIMA(p,d,q)(P,D,Q,s)

  • ss (Seasonality length): The number of time steps in a single seasonal cycle (e.g., s=12s=12 for monthly data with an annual pattern, s=7s=7 for daily data with a weekly pattern).
  • PP (Seasonal AR order): The number of seasonal autoregressive lags. If s=12s=12 and P=1P=1, the model uses the value from exactly one year ago (t−12t-12) to predict today.
  • DD (Seasonal Integration): The degree of seasonal differencing. A D=1D=1 means subtracting the value from one season ago (yt−yt−sy_t - y_{t-s}) to achieve stationarity.
  • QQ (Seasonal MA order): The number of seasonal moving average terms (past seasonal forecast errors).

When it is needed: You need SARIMA whenever exploratory data analysis (like seasonal decomposition or ACF plots showing spikes at regular intervals) indicates strong seasonal behavior. If you attempt to fit a non-seasonal ARIMA model to seasonal data, the residuals will not be white noise—they will contain obvious repeating patterns, meaning the model is leaving valuable predictive information on the table.

💡 Note: SARIMA handles single seasonality well. If your data has multiple seasonalities (e.g., daily patterns and yearly patterns), you may need more advanced models like TBATS, Prophet, or Fourier terms with SARIMAX.

“What is stationarity, and why does it matter?”

Quick answer

A stationary time series has statistical properties—like mean and variance—that remain constant over time. It matters because most traditional forecasting models, like ARIMA, assume stationarity to make reliable predictions.

Answer

Stationarity is a fundamental concept in time series analysis describing a series whose statistical properties do not change over time. Specifically, strict stationarity requires the joint distribution of any sequence of observations to be invariant to shifts in time. In practice, we usually look for weak (or covariance) stationarity, which requires three conditions:

  1. The mean of the series is constant over time.
  2. The variance of the series is constant over time (homoscedasticity).
  3. The autocovariance (and autocorrelation) between two observations depends only on the time lag between them, not on the actual time tt they were observed.

A time series with a strong trend or clear seasonality is not stationary because the mean and variance change depending on when you sample the data.

Why it matters: Most traditional statistical forecasting models, such as ARIMA (AutoRegressive Integrated Moving Average), are predicated on the assumption of stationarity. If the statistical properties of a series are constantly changing, it is mathematically impossible to model it reliably using historical data because the "rules" of the data generating process are shifting.

By applying transformations like differencing (subtracting the previous value from the current value) or taking logarithms, analysts convert a non-stationary series into a stationary one. Once a stationary model is fitted, the transformations can be reversed to produce forecasts in the original scale.

💡 Note: While classical models require stationarity, modern machine learning approaches (like gradient boosting and deep learning) can sometimes model non-stationary data directly if appropriate temporal features and trends are included.

“Why can't you use random train/test splits for time series?”

Quick answer

Random splitting destroys the chronological order and temporal dependencies of the data, leading to data leakage where the model 'sees' future information during training, resulting in overly optimistic performance estimates.

Answer

In traditional machine learning with tabular data, evaluating a model's performance relies heavily on random k-fold cross-validation or random train/test splits. This assumes that all observations are independent and identically distributed (i.i.d.).

In time series forecasting, the data is inherently sequential, and observations are dependent on previous observations. If you use a random split, you will encounter two major, fatal issues:

  1. Data Leakage (Look-ahead Bias): Random sampling means that data from the "future" will end up in your training set, while data from the "past" ends up in your test set. If you are trying to predict the stock price on Wednesday, and your model was trained using the data from Thursday and Friday, it has effectively peeked into the future. It will learn patterns that rely on future information, rendering it useless in a real-world scenario where the future is strictly unknown.
  2. Destruction of Autocorrelation: Time series models rely on features like lags and rolling windows. Randomly plucking rows out of the dataset destroys the continuity required to calculate these temporal features correctly.

To properly evaluate a time series model, you must respect the chronological order. The training set must consist of historical data occurring strictly prior to the data in the test set. Techniques like "Out-of-Time Validation" or "Walk-Forward (Rolling Origin) Validation" are the correct approaches.

💡 Note: An exception exists when dealing with multiple independent time series (e.g., predicting the heartbeat of different patients). You can randomly split by patient ID, ensuring the entirety of one patient's time series is in either train or test, but you still cannot randomly split the timestamps within a single patient's series.