The buying decision was rational. That is what makes this worth writing down.
A retailer looked at several years of sales data, saw one size outselling the others by a wide margin across their range, and bought accordingly. Not a rash decision, not a hunch. They did the thing everyone says you should do. They let the numbers lead, and the numbers pointed at that size consistently, year after year, across stores.
The stores were not scanning.
Rather than passing a barcode over a scanner, staff were typing the product code into the till and picking a size from the list that came up. Typing a code is faster than finding a barcode on a shoebox, especially at a busy till, and the list that appears is sorted. So they picked the first one. Thousands of times a day, across a chain, for years.
Every one of those transactions was recorded correctly. The till captured exactly what was entered. The sales data was accurate in the only sense a database understands. It was also describing the sort order of a dropdown menu rather than the feet of any actual customer, and the business had put millions of rands into stock on the strength of it.
Accurate and true are different properties
Data quality work almost always focuses on accuracy. Does the number in the system match the transaction that happened? Here, it did. Every record was faithful to the event that produced it, and any validation you care to run would have passed.
The failure was one layer further back, in what the event actually represented. The system recorded a keystroke. Everyone downstream read it as a customer preference. Nothing in the data marks the difference, and no amount of downstream rigour recovers it.
This is the uncomfortable part, because it means the usual defences do not apply. Reconciling would not have caught it, since the totals were right. Auditing the report would not have caught it, since the report faithfully summarised the records. More dashboards would have made it worse by giving the same broken signal more surfaces to be seen on. The only thing that catches it is somebody asking what the number is physically a record of, and that question gets asked roughly never once a report has been running for a while.
How to test your own sales data in about ten minutes
Four checks. None needs a project, and any of them can be done in whatever you already have.
Compare a size's sales rank to its position in the list. Pull the best-selling size for twenty unrelated styles. If the winner is consistently the one that sorts first, you are not looking at a customer preference. Real demand does not care about alphabetical order. This is the check that would have caught the case above in an afternoon.
Compare a scanning store to a non-scanning one. If two branches with similar customers report meaningfully different size curves for the same product, the difference is in the till, not the customers.
Compare what sells to what runs out. This is the strongest signal and the one people most often already know at gut level. If you are constantly out of stock in a size the data says is slow, and sitting on a size the data says flies out, the data is describing something other than sales. Your store managers have probably been saying this for a while.
Compare the size mix of sales to the size mix of receipts. If they track each other suspiciously closely, over a long period, be suspicious. Sales should diverge from purchasing, because purchasing is a guess and sales are the answer. When the answer looks exactly like the guess, something is echoing.
It is a process problem wearing a data costume
The instinct is to fix this in software. Force scanning, block manual entry, validate the size against something.
That instinct is half right and it usually fails, because the reason staff were typing codes was that typing codes was faster at a busy till, and a rule that makes the till slower during a queue gets worked around within a week. If manual entry is blocked, staff will find the fastest legal path, and the fastest legal path will produce its own systematic bias that nobody has thought about yet.
What actually works is duller. Make scanning the fast path rather than the compliant one, which usually means the scanner placement and the barcode position on the packaging, not the software. Then measure entry method as a first-class number: what percentage of lines were scanned, by store, by hour. The moment that percentage is visible on a report somebody looks at, it starts improving, and you also gain the ability to say how much of your history you can trust.
That last part is the real deliverable. Not clean data going forward, which takes a season or two. The ability to put a confidence level on the data you already have.
- Sales data can be completely accurate and still be recording something other than what customers bought.
- If the same size wins across unrelated styles, check whether it is simply the first in the list.
- Selling well and running out should broadly agree. Where they disagree, trust the stockouts.
- Track the percentage of lines scanned versus typed. It tells you how much of your history you can rely on.
What it actually costs
The obvious cost is the stock. In the case above it ran to millions of rands committed to a size that was never selling in the volumes the reports claimed. Buy too deep into one size and you carry it, mark it down, and eventually move it at a loss.
The larger cost is the one nobody books. Every unit bought in the wrong size was a unit not bought in a right one, so the loss is the markdown plus the sales that never happened because the size somebody wanted was not there. That second number is invisible, it does not appear in any report, and it is usually the bigger of the two.
There is a third cost that is worth naming even though you cannot price it. Once buyers learn that the sales data misled them, they discount it, and they go back to buying on judgement. Judgement is not worthless, but it is not better than good data either, and a business that has quietly stopped trusting its own numbers is in a worse position than one that never had them.
Where this does not apply
If your stores genuinely scan everything, this is not your problem, and the four checks above will come back clean in a morning. Run them anyway once, because a clean result is worth knowing with certainty rather than assuming.
The checks also only tell you the recorded curve is wrong. They will not tell you the true one. Recovering that takes a clean season of properly captured data, and in the meantime you are making decisions with a known unknown, which is uncomfortable but is at least honest, and is a great deal better than confident and wrong.
The question worth asking
Not "is this number accurate", which it usually is.
"What is this number physically a record of?"
For a size curve, the honest answer might be "a keystroke by someone who was in a hurry". That is still useful information. It is just not the information everyone thought they were buying against.
