Customer data quality: don't let inaccurate data stop you from acting
Customer data quality is never perfect. How to judge whether data is good enough for the decision at hand, triage the flaws, and act without fooling yourself.
Table of contents
- Key takeaways
- What customer data quality actually means
- What inaccurate customer data looks like in practice
- Why waiting for clean data costs more than acting on imperfect data
- How to triage flawed customer data before a decision
- The caveat box: the habit that makes imperfect data safe
- When customer data quality really does block action
- Where to start
- FAQ
The monthly voice-of-customer review has reached the slide about billing. Complaints about invoices are up. The comments are specific and unhappy. Then someone asks about the response rate, and the answer is 12 percent, and the room relaxes. Probably just the angry ones. Let’s wait until we have better data. Customer data quality has just been used as a reason to do nothing.
Customer data quality is the degree to which your customer records, survey responses and behavioral data are fit for the decision you are about to make with them. Fitness is the key word: the same dataset can be good enough for one decision and useless for another.
I have watched some version of that billing scene many times, and I have come to think of it as a kind of procrastination that looks like rigor. Nobody in the room is being lazy. They are being careful. But the effect is that a signal was received and nothing happened, and next month the same slide will be back with a slightly larger number on it.
Key takeaways
- Customer data is never clean, so the useful question is whether it is accurate enough for the specific decision in front of you, not whether it is accurate.
- Most customer experience decisions are “should we look at this or not,” and for those, direction is enough; precision matters only for fine-grained calls such as small price changes.
- Flaws sort into three piles: fix the ones that block the next step, flag the ones that could mislead, and ignore the ones that do not touch this decision.
- A caveat box under every number, stating source, gaps and what would change the conclusion, is what makes acting on imperfect data safe.
- A skewed survey sample that over-represents unhappy customers is a fine early-warning system, because the unhappy ones are the ones about to leave.
What customer data quality actually means
Data quality is usually described along a handful of dimensions, and it helps to name them, because “inaccurate” hides which one is broken.
Accuracy is whether a value is right: the email address works, the purchase amount matches the invoice. Completeness is whether the value is there at all: the postal code field is empty for a third of records. Timeliness is whether it is current: the contact left the company two years ago. Consistency is whether two systems agree: the customer is “active” in one and “lapsed” in another. Linkage is whether records about the same person can be joined: the survey response cannot be matched to an account.
Each dimension fails differently, and each failure matters for different decisions. A stale postal address is a problem for a mailing and irrelevant to a complaint analysis. A survey that cannot be linked to accounts is a problem for follow-up calls and irrelevant to spotting which topic is rising.
The definition that matters in practice is fitness for purpose: good enough to act on, given what the decision costs if it is wrong.
What inaccurate customer data looks like in practice
The flaws are not all the same flaw, and they live in different places.
The customer database has duplicates: the same person under two email addresses, or a household counted as three accounts. A share of the email addresses bounce, and some of the ones that do not bounce belong to people who left the company two years ago. Purchase history is missing the transactions from the system that was replaced in a migration. Why can’t we collect email addresses? is about how one of those gaps opens up and why it is so hard to close.
Survey data has its own problems. The people who respond skew toward the ones who are delighted and the ones who are furious. The middle is underrepresented, and the middle is where most of your customers live.
And behavioral data, which is supposed to be the objective one, has its own holes: sessions that did not get tracked, events that fire twice, a definition of “active” that changed in March without anyone updating the report.
| Data source | Typical flaw | Decisions it affects | Decisions it does not |
|---|---|---|---|
| Customer database | Duplicates, dead contacts, migration gaps | Mailings, contact-level follow-up, counting customers | Which complaint topic is rising this month |
| Survey responses | Low response rate, skew toward the extremes, unlinked to accounts | Estimating what share of all customers feel something | Whether a topic is rising, what the specific complaint is |
| Behavioral data | Untracked sessions, duplicate events, changed definitions | Precise conversion or usage rates, trend lines across a definition change | Whether a specific group did the core thing at all |
All of that is true, and all of it has been true at every company I have worked with. That is the normal condition of customer data. A cleanup project will improve it and will not end it.
Why waiting for clean data costs more than acting on imperfect data
The question to ask about any piece of customer data is not “is this accurate?” It is “is this accurate enough for the decision in front of me?”
A survey with a 12 percent response rate cannot tell you that some exact share of customers is unhappy about billing. It can tell you that billing is the thing people are complaining about most this month, and that the complaints are more specific and more heated than last month. That is enough to open the invoices and look. It is enough to call ten of the respondents. It is enough to put someone on it for a week.
The skew toward the angry actually helps here. Angry customers are the ones about to leave. A sample that over-represents them is a sample that over-represents the risk, which is exactly what an early-warning system should do. The fire alarm does not need a representative sample of the building’s air. It needs to go off when there is smoke. The customers who leave without complaining, the subject of You’re fired, are the ones you will never hear from, so the ones you do hear from deserve attention out of proportion to their number.
Precision matters when the decision is fine-grained: whether to move a price by two percent, whether one variant of a page beats another by a hair. Most CX decisions are not like that. Most of them are “should we look at this or not,” and for that, direction is enough.
The cost of waiting is easy to miss because it never appears on a slide. Nobody records the customers who left during the quarter spent improving the data. The cleanup project has a budget line; the delay does not.
How to triage flawed customer data before a decision
When the data is flawed and the decision cannot wait, work through these steps. They take an hour, not a quarter.
- Name the decision. Write one sentence: “We will decide whether to put someone on the billing complaints for a week.” Everything that follows is judged against that sentence, not against an abstract standard of accuracy.
- List the known flaws. Response rate, skew, missing channels, duplicates, the definition that changed in March. Most teams find the list is shorter than the anxiety suggested.
- Sort each flaw into one of three piles. Fix, flag or ignore, as described below.
- Fix only the first pile, crudely if necessary. The aim is to unblock the next step, not to produce a clean dataset.
- Act, with the second pile written next to the number. Then note what you did and what you saw, so next month’s slide has a comparison.
The three piles
Fix what blocks action. If you cannot tell which customers complained about billing because the survey was anonymous and unlinked, that blocks the follow-up call. Fix that, even crudely: match on order number, ask the next batch of respondents for permission to contact them. If duplicate records mean you would email the same person three times, dedupe before sending. Only the flaws that stop you doing the specific next step belong in this pile.
Flag what could mislead. Some flaws do not stop you acting but could send you the wrong way. The survey over-represents recent buyers because that is who got the link. The behavioral data is missing one channel. These get written down, in plain words, next to the chart, so that everyone who sees the number sees the caveat. Then you act anyway.
Ignore what does not matter for this decision. The billing complaints are the same whether or not the customer’s postal address is current. The migration gap in purchase history from four years ago does not change what happened on last month’s invoices. Naming the flaws that are irrelevant is as important as naming the ones that count, because otherwise every flaw becomes a reason to wait.
That last pile is where data perfectionism lives. A team that cannot say “this problem does not affect this decision” will find a reason to postpone every decision.
A worked example (illustrative)
Take the billing slide. The decision: whether to assign one person for one week to read the invoice complaints and call ten respondents. Known flaws: 12 percent response rate, skew toward unhappy respondents, survey link sent only to customers who bought in the last 90 days, a fifth of responses unlinked to an account, and a purchase-history gap from a migration four years ago.
Sorting: the unlinked responses partly block the calls, so fix by matching on order number where possible and calling from the linked four-fifths. The response rate and the skew could mislead a claim about how widespread the problem is, so flag them and make no such claim. The 90-day sampling window is flagged too, since older customers may see different invoices. The migration gap touches nothing here, so ignore it. Total time to triage: about an hour. Time the same team had been waiting for better data: two months.
The caveat box: the habit that makes imperfect data safe
The one practice that makes it possible to act on imperfect data without fooling yourself is the caveat box. Every chart, every table, every number that goes into a deck gets two or three lines underneath it: where the data came from, what is missing, and what would change the conclusion.
“Survey of customers who purchased in the last 90 days. Response rate 12 percent. Respondents skew toward recent and higher-frequency buyers. If the same pattern does not appear in support tickets, treat with caution.”
Written that way, the number can be trusted for what it is and not for what it is not. It also builds the kind of confidence that sharing data across teams depends on. People stop fearing that a number will be used against them once the gaps travel with it. And it makes room for the qualitative side, the comments and the calls, which Bringing a soft focus to hard data argues is what tells you what the numbers mean.
The third line, what would change the conclusion, is the one most teams skip and the one that matters most. It turns a caveat from a disclaimer into a test. If the support tickets do not show the same billing pattern, the survey signal was probably noise, and you have learned that in a week instead of never.
When customer data quality really does block action
Fitness for purpose cuts both ways. Some decisions genuinely need better data than you have, and pretending otherwise is the opposite error.
Contact-level action on unlinked data. If the next step is to call, email or credit a specific customer, you need to know who they are. Anonymous feedback can point at a problem but cannot tell you whom to call about it.
Anything with legal or consent weight. Sending marketing to addresses whose consent status is unknown, or acting on personal data whose provenance is unclear, is a compliance question before it is a data quality question. Fix the record first.
Fine-grained financial decisions. Repricing, changing a fee, or reallocating a large budget on the strength of a small skewed sample is how bad data does real damage. These need a representative sample or a controlled comparison.
Trend claims across a definition change. If “active” was redefined in March, a chart from January to June is two charts stapled together. Either restate the earlier months or split the line.
For everything else, which is most of what a customer experience team decides, the triage is enough. Whoever owns the customer analysis should be the one saying, in each case, which side of the line the decision falls on.
Where to start
- Find the last customer signal your team decided to wait on. There is usually one within the last quarter.
- Write the decision it was supposed to inform in one sentence. If nobody can, that is a different problem, and a bigger one than data quality.
- List the flaws and sort them into fix, flag and ignore. Do it in a meeting with the people who raised the doubts, so the ignore pile is agreed rather than imposed.
- Write the caveat box. Source, gaps, and what would change the conclusion, in three lines under the chart.
- Act on the signal in the smallest way that would teach you something. Ten calls, one week of one person’s time, one look at the actual invoices.
- Bring the result back next month, next to the caveat box. Whether the signal held up or not, the team has now practiced acting on imperfect data once, and the second time is easier.
FAQ
What is customer data quality?
Customer data quality is how fit your customer data is for the decision you intend to make with it. It is usually assessed along dimensions such as accuracy, completeness, timeliness, consistency between systems, and whether records about the same person can be linked. The same dataset can be high quality for one purpose and inadequate for another.
How do you assess whether customer data is good enough to act on?
Start by writing down the specific decision, then list the known flaws in the data and ask of each one whether it blocks the next step, could mislead the conclusion, or does not touch the decision at all. Fix the first kind, write the second kind next to the number, and ignore the third. Data is good enough when acting on it is more likely to be right than waiting, given what a wrong decision would cost.
Is a low survey response rate too low to act on?
A low response rate limits what you can claim about the whole customer base, but it rarely prevents you from acting on direction. A survey answered by a small share of customers can still show which topic is rising and what the specific complaints are, which is enough to investigate. It is not enough to state what percentage of all customers feel a certain way.
What are the most common customer data quality problems?
The most common are duplicate records, contacts that are out of date, transaction history lost in system migrations, survey responses that cannot be linked to an account, samples that over-represent the very happy and the very unhappy, and behavioral definitions that change without the reports being updated. Every company has some of these at all times. The useful skill is knowing which of them matter for the decision at hand.
Should you clean customer data before starting a voice-of-customer program?
No. Fix only what blocks the first actions, such as being able to reach the customers who respond, and start listening. A cleanup that has to finish before anything happens usually delays the program by months and never finishes. The program itself surfaces the data problems that actually matter.
How do you present imperfect customer data to executives?
Put a short caveat under every number stating where it came from, what is missing and what evidence would change the conclusion. Present the finding as a direction rather than a precise share, and pair it with the small action you propose to take. Executives are used to acting on incomplete information; what they distrust is a number whose gaps were hidden.