The signal is the truth. The noise is what distracts us from the truth. – Nate Silver

Do you know (in PSI or kPa) what a proper oil pressure reading is for your car? Or the correct voltage reading for your alternator? Or at what temperature your engine should be running?
I sure don’t and I’d wager that the vast majority of car owners don’t either. That’s why over time, measurement gauges on a car’s dashboard have been replaced (or augmented with) warning lights – something to let you know that there is a problem that requires your attention. (As an aside, my grandfather who was a mechanic for 50+ years used to refer to them as “idiot lights” – as if anyone who doesn’t readily know the answer to my questions above shouldn’t even be allowed to drive).
While it’s true that gauges provide more information (a range of possible values versus just a binary on/off signal), you really need to fully understand the capabilities and limits of the systems being measured to know what they’re actually telling you and – more importantly – if action is required based on the reading, which you would need to monitor constantly.
I have been in the supply chain planning game for over 30 years now and when I meet a new colleague for the first time, it’s not uncommon for him/her to lead with something like “Our forecast accuracy is 79.6%” (I suppose they must be looking for some sort of affirmation or praise).
What they end up getting is usually a confused look and about a thousand questions:
- Are you measuring at item/location level then aggregating or aggregating then measuring?
- What specific measurement are you using?
- Over what time frame?
- Are the errors biased in one direction or the other?
- Etc., etc., etc.
However, the most important questions (which take some time to come to the surface) are:
- How do you know if that’s good or bad?; and
- What do you do with that information?
When you multiply the number of items through the number of locations for a retailer, there are a lot of forecasts being produced across fast sellers, slow sellers, seasonal items, highly promoted items, products at the beginning/end of their lifecycle, not to mention myriad combinations of all of those attributes.
There’s nothing mathematically invalid about aggregating numbers and dividing them by each other to come up with a percentage using the full data set, but there’s nothing intrinsically useful about it either. Unless – and this will sound a bit cynical – your goal is to produce a number that makes it appear as though the process is working beautifully, even though it may be completely skewed by very high volume and stable selling items that require no effort to forecast well.
Most importantly, it doesn’t give a true picture of the highly varied landscape that is retail forecasting and it’s not actionable.
The forecasts that matter are the ones that directly affect the customer experience – for each specific item at each specific location. I’ve discussed this before, but when you’re looking at the individual granular forecasts that actually run the supply chain, you need to be accurate, not precise (those are two different things). It means moving away from comparing a single forecast to a single actual and calculating percentages to determine distance from absolute perfection (a completely unreasonable yardstick to measure against).
It’s really about recognizing that for any given item at any given store at any given time, there is a range expected forecast values that can be considered high quality, depending on the nature of the product, the timing of the demand and the additional causals that might be in play at the time.
Think of it this way: If your forecast for an item at a store is 4 units for the week and you actually sold 2, your “accuracy” is 50%, which sounds bad. But is it really a bad forecast? Is there something you could have done in advance that would have led you to forecast a value closer to 4? Now that you know you missed by 2 units, what specific actions can you take to make sure the next forecast is closer? What you end up doing in this scenario is chasing noise – and in retail forecasting there is a lot of it.
However, at this level of volume, it’s probably quite common to miss by 3 units in either direction in any given week, just as a result of uncontrollable randomness. In other words, there is an expected and acceptable “distance from perfect” that applies to every forecast. To make your forecast measurement truly actionable and useful, you should be measuring actuals against this range to make this determination. So, in our example above, if the actual is 4 and the common range of error is +/- 3 units, then any forecast between 1 and 7 units would be considered accurate when compared against the actual.
If you use a probabilistic forecasting approach, these acceptable ranges can be calculated with the forecast itself. If not, there would need to be some upfront analysis to assign one to each forecast, but the end result is that you have a way to find the needles in the haystack: instead of knowing which forecasts have error (pretty much all of them), you instead know which forecasts have unexpectedly high error. That’s actionable information that can be chased down and investigated.
With each individual forecast classified in this way (essentially an accurate/inaccurate “warning light”), you can also present a more meaningful gauge on your high level dashboard. Instead of a number representing “distance from unattainable perfection in the aggregate”, you can instead report on percentage of forecasts out of error tolerance.










