The question “is our forecast good enough?” almost always comes down to another one: was the method matched to the data actually available? Forecasting demand based on thirty monthly observations and on five years of weekly transaction data are two different analytical tasks, even if the business decision may be identical. Below is an overview of how to choose between exponential smoothing and an ARIMA model – and how to verify that the choice was appropriate.
When does demand forecasting require a simple model, and when a complex one?
The starting point is not analytical ambition, but an inventory of the data. Demand forecasting is based on a time series, and the choice of methods depends primarily on the length of the historical data, observation frequency, seasonality, data quality, and noise level.
Length of historical data. Methods that require estimating multiple parameters need a sufficiently long series. With a dozen to twenty-something monthly observations, it is realistically possible to reliably estimate the level and – cautiously – the trend. Attempting to fit an extensive autoregressive model to such a short series may lead to overfitting: the model describes the past well but makes errors in the future. If the series contains a pattern that repeats on an annual cycle, at least two to three full cycles are needed for reliable estimation, and weekly data usually require a correspondingly longer history.
Observation frequency. Annual data provide few data points and usually do not allow for anything more than a level model with trend. Monthly data are standard ground for both method families. Weekly and daily data introduce additional structure (calendar effects, day-of-week effects, movable holidays), which classical exponential smoothing will not account for unless it is extended with a regression model, while ARIMA will only account for it when extended with external variables.
Noise level and the nature of demand. A smooth series with a clear direction and regular pattern is a natural environment for smoothing methods. A series with strong residual autocorrelation, inertia in response to previous values, and time lags is an environment for the ARIMA class. Intermittent demand, with a large number of zero periods, typical of spare parts and low-turnover products, is a separate category; in such cases, neither method is often the right tool, and approaches specifically designed for intermittent demand are used.
In practice, the decision on the method is made after diagnostics, not before. The analysis includes time series decomposition, stationarity testing, examining the autocorrelation and partial autocorrelation functions, and identifying outliers resulting from promotional campaigns, distribution changes, or one-off events. Only then does the evidence indicate whether sales forecasting in a given case requires a model that better describes autocorrelation, or whether a mechanism that adaptively updates the level and trend is sufficient.
How to choose a demand forecasting method in practice?
Exponential smoothing is a family of methods in which the forecast is a weighted average of past observations, with weights decreasing exponentially as observations recede into the past. The simplest version models the level only. The version with trend (Holt) adds a directional component. The Holt-Winters version includes level, trend, and a recurring seasonal pattern component. The modern formulation of this family is ETS state space models, in which the combination of components and their type (additive or multiplicative) is selected based on information criteria.
The ARIMA model works differently: it describes a series through the relationship between the current value and lagged values (the autoregressive component), past forecast errors (the moving average component), and differencing, which can remove nonstationarity. The SARIMA variant adds analogous components at the seasonal level, while the variant with regressors (ARIMAX, SARIMAX) allows external factors to be included: price, promotional intensity, number of points of sale, temperature, and the holiday calendar.
Practical selection criteria can be arranged in sequence. Below is a set of questions that most often determine the choice in forecasting projects:
- How many observations are available? A short series – exponential smoothing or a naive model with adjustment. A long series – opens up scope for ARIMA.
- Are the residuals of a simple model random? If clear autocorrelation structure remains in the residuals after fitting exponential smoothing, the simple model has not exhausted the information contained in the data, and the ARIMA class may have more to work with.
- Do explanatory variables need to be included? Classical exponential smoothing does not accept regressors. If the forecast is expected to respond to planned price reductions or distribution changes, a model with a regression component is needed.
- What is the forecast horizon? Over a short horizon, differences in accuracy between methods often prove small. Over a longer horizon, the correct specification of the trend and its damping becomes more important.
- How many series are being forecast? With hundreds of SKUs, automation and stability matter. The ETS family is easier to use at scale in an automated way; automatic ARIMA identification is also possible, but requires stricter result control.
- Who will maintain the model? A model that the team does not understand and cannot rerun stops being used after a few months.
However, the decision is not made at the diagnostic stage, but at the validation stage. The standard approach is to split the series into training and test sets, although – unlike with cross-sectional data – the split must respect the temporal order. A more demanding and informative approach is rolling validation: the model is repeatedly estimated on an increasingly long historical window and evaluated on subsequent periods it has not seen. This produces a distribution of errors rather than a single random result.
As Hume’s Institute experts point out, a more complex model does not guarantee a better forecast – the choice is determined by testing on data the model has not seen. In-sample fit usually improves as parameters are added, so comparing methods solely on this basis favors the selection of overfitted models.
The comparison must be anchored to a benchmark. In time series forecasting, this role is fulfilled by naive models: a forecast equal to the most recent observation, a forecast equal to the value from the corresponding period of the previous season, or a moving average. If neither exponential smoothing nor ARIMA improves on a seasonal naive model, this means they do not use the available structure better than a simple benchmark. It is then worth checking the model specifications, data quality, and the possibility of including relevant external variables. The choice of error metric and interpretation of its values is a separate issue, discussed in more detail in the material on what MAPE and other error metrics really indicate.
What errors most often undermine demand forecasting?
Most inaccurate forecasts do not result from choosing the wrong model family, but from errors made earlier – in the data – or from an incorrect evaluation procedure. Below are the pitfalls most commonly encountered in forecasting projects.
- Confusing sales with demand. Sales data may be constrained by product availability. A period of out-of-stock conditions may be recorded as low demand, and the model learns a false decline. Without adjusting for inventory levels and availability, demand forecasting may turn into forecasting supply constraints.
- Unlabeled one-off events. A promotional campaign, listing with a new retailer, a one-off wholesale order, or a packaging change – each of these events disrupts the series. If they are not labeled as regressors or outliers, the model may spread their effect across the entire history.
- Evaluating the model on training data. One of the most common methodological errors. A model with more parameters usually achieves a better fit in the sample on which it was estimated, which does not determine its quality beyond that sample.
- Omitting a benchmark model. Without comparison against a naive forecast, it is impossible to say whether the model adds anything. Information that the error is a few percent does not in itself say anything about the quality of the method.
- Differencing by intuition. In the ARIMA class, excessive differencing increases forecast variance and may worsen performance over a longer horizon. The degree of integration should result from series diagnostics, not habit.
- A point forecast without an interval. A single number suggests a certainty that does not exist. The forecast interval is part of the result, not an add-on, and its width indicates how predictable the series is.
- Failure to refresh the model. The structure of the series changes with the market. A model estimated once and used for years gradually loses accuracy, and the re-estimation process should be planned from the outset.
- Forecasting at the wrong level of aggregation. Individual SKU series are usually noisier than category series. A forecast at an aggregated level and its allocation to lower-level items based on shares often produces more stable results than modeling each item independently.
It is also worth honestly acknowledging the limitations of both method families. Exponential smoothing and ARIMA are forecasting methods based primarily on the internal structure of the series, while variants with regressors additionally depend on the availability of reliable data on explanatory variables. They will not predict the entry of a new competitor, a regulatory change, or a sudden shift in consumer behavior if information about these events is not introduced into the model. Nor do they forecast products without historical data – in such cases, analogous data, purchase intention research, and concept testing models are the foundation, while time series forecasting comes into play only after the first sales periods have been collected.
How do typical use cases of the two methods differ?
Below is a summary of situations in which the choice is usually clear. This is not an exception-free rule, but a starting point for validation.
- Short history, monthly or quarterly data, one product. Exponential smoothing with a damped trend. An extensive ARIMA model may be too unstable for reliable estimation.
- Several years of monthly data, a clear recurring annual pattern, no external variables. Holt-Winters or ETS as the first choice, with SARIMA as a competing candidate. Rolling validation determines the choice.
- Several years of weekly data, intensive promotional activity. SARIMAX with regressors describing promotions and the calendar. Classical smoothing methods will not use this information unless extended with a regression component.
- Hundreds of SKUs, forecast refreshed monthly, limited analytical team. Automated ETS with quality control and a naive model as the benchmark.
- A series with strong inertia, delayed response, and dependence on prior deviations. The ARIMA class is a natural candidate because it describes this type of structure.
- Intermittent demand, many zero periods. Methods designed for intermittent demand, not ETS or ARIMA.
In Hume’s Institute projects, the most useful demand forecasts are observed when, rather than selecting a single “best” model, several candidates are tested using a common validation framework, and the result is not only a number but also a description of uncertainty and a list of factors not covered by the model.
Frequently asked questions
What are the methods of demand forecasting?
They fall into three groups: quantitative methods based on time series (naive models, moving averages, exponential smoothing and the ETS family, ARIMA and SARIMA, and models designed for intermittent demand), causal methods that link demand to explanatory variables (regression, ARIMAX, machine learning models), and qualitative methods used when historical data are unavailable (purchase intention research, concept tests, expert panels, and the Delphi method). In practice, these are combined: the quantitative forecast provides the baseline, while qualitative knowledge adjusts it for events absent from the data.
How does the ARIMA model differ from exponential smoothing?
Exponential smoothing builds a forecast as a weighted average of past observations, with weights declining exponentially, and explicitly models the level, trend, and recurring seasonal pattern components. The ARIMA model describes a series through dependence on lagged values, past forecast errors, and differencing, which can remove nonstationarity. In practical terms, ARIMA may handle autocorrelation better, and its variants with regressors accept external variables. Exponential smoothing, on the other hand, is often easier to apply automatically across many series.
How much historical data is needed for a demand forecast?
The minimum depends on which components are to be modeled, the data frequency, and the forecast horizon. For forecasting level and trend using monthly data, around two years of observations may be a reasonable starting point, although a shorter history may also be sufficient in simple cases. If the series contains a recurring annual pattern, at least two to three full cycles are needed, and substantially more data are needed for stable SARIMA estimation using weekly data. Beyond length, completeness matters: gaps, changes in measurement definitions, and unlabeled one-off events reduce the usefulness of historical data more than its length does.
Ask about a demand forecast for your category
If an assessment is needed of whether the available data allow for a reliable forecast and which method is appropriate, Hume’s Institute will prepare a series diagnostic and compare models using external validation. Simply contact Hume’s Institute with a brief description of the data and forecast horizon.