Forecast accuracy: what MAPE and other error measures really tell you

Monika

Your forecasting model shows a MAPE of 8%, which looks impressive – until it turns out that errors in a key product category reach several dozen percent and are simply obscured by averaging. Assessing forecast accuracy using MAPE can be misleading if it is treated as the only number in the report. This article explains what different forecast error measures actually capture and how to combine them so that the assessment of forecast accuracy is fair rather than merely cosmetic.

What does forecast accuracy measured by MAPE really tell you?

MAPE (Mean Absolute Percentage Error) is the average absolute percentage error – the sum of deviations between the forecast and actual values, expressed as percentages and averaged across all periods. Its popularity stems from its simple interpretation: the result is reported as a percentage, so a manager without a statistical background understands that “on average, we are off by this many percent.” This is a communication advantage, but it is also the source of most misunderstandings surrounding MAPE forecast accuracy.

The problem is that MAPE is not a neutral measure. It divides the absolute error by the actual value, so each low-volume period carries disproportionately high weight in the result. An error of a few units when sales are in the low double digits generates a huge percentage error, whereas the same deviation at high volumes barely affects the metric. In practice, this means that MAPE penalizes a model for inaccuracy where the business stakes are lowest and rewards it for apparent precision where the numbers are large.

The second pitfall is asymmetry. MAPE treats forecast overestimation differently from underestimation. Because the denominator is the actual value, overestimated forecasts can theoretically generate an unlimited percentage error, whereas underestimation, with non-negative forecasts, is mathematically capped at 100%. A model that consistently underestimates may perform better on MAPE than a more ambitious model, even if both make similarly sized errors in absolute terms. This is not a minor nuance – it is a systematic distortion of the assessment.

The third issue is that MAPE does not work at all when zeros appear in the data. Division by zero is undefined, so every period with no sales – typical of seasonal products, new launches, or items with intermittent demand – must either be excluded or handled using workarounds that distort the result. This is why forecast accuracy assessed solely through MAPE can be useless precisely where forecasting is most difficult.

How should MAPE, WAPE, and RMSE be combined to ensure a fair assessment of forecast accuracy?

No single measure fully describes forecast accuracy, because each answers a different question. A meaningful assessment of forecast accuracy involves reading several metrics in parallel and drawing conclusions from the differences between them. Below are three measures that, combined, provide a picture of where the model is actually making errors:

  • MAPE – answers the question “by what percentage are we wrong on average,” but weights periods inversely to their volume and breaks down when zeros are present. It is useful for communication but risky as the sole criterion.
  • WAPE (Weighted Absolute Percentage Error, also called MAD/Mean) – sums absolute errors and divides them by the sum of actual values. As a result, it weights errors in proportion to volume, handles individual zeros as long as the sum of actual values in the analyzed dataset is not zero, and better reflects demand forecast error in categories with varying volumes.
  • RMSE (Root Mean Squared Error) – the square root of the mean of squared errors, expressed in the original units rather than as a percentage. Squaring the errors makes RMSE highly responsive to large, isolated errors. It is a measure sensitive to outliers.

The greatest diagnostic value lies in comparing these figures. When MAPE is high and WAPE is low, the signal is clear: the model is making errors mainly for low-volume items while forecasting high-volume items accurately. The reverse pattern suggests that the problem may concern key, high-volume products. A gap between linear measures and RMSE, in turn, indicates the presence of infrequent but very large errors – RMSE rises much faster than linear measures when the data contain a few major misses.

As Hume’s Institute experts point out, a single error figure never tells you whether a forecast is good – only a combination of several measures shows where the model is actually making errors and whether those errors are systematic or random. A report with one MAPE value is a report that conceals more than it reveals.

It is also worth adding one more measure to this set, one that concerns not the size of the error but its direction. Bias, or the signed mean error, shows whether the model systematically overestimates or underestimates forecasts. Absolute measures such as MAPE or WAPE lose this information because they treat upward and downward deviations in the same way. Yet systematic underestimation of demand forecasts leads to different operational consequences than systematic overestimation, and no absolute error measure will capture this.

In practice, the combination of MAPE, WAPE, RMSE, and bias creates a set that answers four different questions: how large the relative error is, how large the volume-weighted error is, whether extreme misses occur, and whether the model tends in one direction. Analyzing these four dimensions together turns the assessment of forecast accuracy from a reporting ritual into a diagnostic tool.

When do forecast error measures lead you astray?

The most common methodological mistake is comparing MAPE across different datasets without considering their characteristics. It is safest to compare the MAPE of two models when they forecast the same time series or series with a similar volume structure and level of variability. Comparing MAPE forecast accuracy for a high-volume product with that of a niche product can be misleading – a lower figure for the high-volume product does not necessarily mean a better model; it may simply indicate an easier task.

The second pitfall is aggregation. An average MAPE across an entire portfolio can look healthy while actually masking dramatic discrepancies across individual segments. In Hume’s Institute projects, breaking down forecast error measures by category, channel, or region often reveals where the problem really lies – the aggregate metric frequently looks better than its components. The assessment of forecast accuracy should begin with disaggregation, not with a single figure on a slide.

The third risk is ignoring the benchmark. A MAPE in the low double digits says nothing on its own until it is known what error is generated by a naive model – for example, a forecast assuming that tomorrow will be the same as yesterday, or repeating the value from a year ago. If an advanced model does not outperform a simple benchmark, its low absolute error is meaningless. Relative measures that compare the model’s error with the error of a naive forecast protect against complacency based on a number that merely looks good.

It is also worth considering the forecast horizon. Demand forecast error for the coming week and for a quarter ahead are two different things, and combining them into one metric blurs the picture. Forecast accuracy usually declines as the horizon lengthens, so error measures should be reported separately for each time interval that matters for decision-making.

How should you select an error measure for a specific task?

The choice of measure should follow from the nature of the data and the purpose of the forecast, not from habit. Below are practical criteria that help structure the decision:

  1. When volumes vary and the data contain zeros – choose WAPE rather than MAPE. It weights errors proportionally and does not break down for individual zero-value observations, as long as the sum of actual values in the analyzed dataset is not zero, making it better for assessing demand forecast error across a broad portfolio.
  2. When the cost of a single large error is high – add RMSE. Its sensitivity to outliers is an advantage wherever one major miss is more costly than many small ones.
  3. When the report is intended for a non-technical audience – MAPE can be useful for communication, but never as the only figure. Always report it alongside WAPE and include commentary on the error structure.
  4. When the direction of the error matters – report bias. It is the only one of these measures that will show the model’s systematic tendency to overestimate or underestimate.
  5. Always – compare the result with a naive benchmark. Without a reference point, no figure indicates whether the model is good or merely looks acceptable.

When assessing forecasting models used in market research, the same logic applies whether the forecast concerns sales, inquiry volumes, or shares. A set of several forecast error measures, disaggregated results, and comparison with a benchmark are the minimum methodological requirements that distinguish a fair assessment of forecast accuracy from a report tailored to produce a favorably sounding number.

Frequently asked questions

What is MAPE?

MAPE (Mean Absolute Percentage Error) is the average absolute percentage error – a measure showing by what percentage, on average, a forecast differs from the actual value. It is calculated as the average of absolute deviations expressed as percentages of the actual value. It is valued for its ease of interpretation but is sensitive to low volumes and undefined for zero values.

When is MAPE misleading at low volumes?

MAPE divides the absolute error by the actual value, so at low volumes even a small error in the number of units produces a very large percentage error. As a result, periods with low sales dominate the result even though their business significance is limited. For series containing zeros, MAPE does not work at all – in such situations, WAPE performs better, provided that the sum of actual values in the analyzed dataset is not zero.

Which error measure should be chosen for demand forecasting?

For demand forecasting with varying volumes, WAPE often works well because it weights errors in proportion to sales volume and handles individual zeros, provided that the sum of actual values is not zero. It is worth supplementing it with RMSE if large individual errors are costly, and with bias to detect systematic overestimation or underestimation. None of these measures should be interpreted without reference to a naive benchmark.

Ask about assessing forecast accuracy for your company

If you want to check whether your forecasting models are making errors where they truly matter financially, Hume’s Institute experts can help select the right set of error measures and conduct a reliable assessment of forecast accuracy. Contact Hume’s Institute to discuss the scope of the analysis.