Qualitative and quantitative feedback: a soft focus on hard data
Qualitative and quantitative customer feedback answer different questions: scores say where, comments say why. How to pair them and read 500 verbatims fast.
Table of contents
- Key takeaways
- What qualitative and quantitative customer feedback are
- Why scores alone lead meetings astray
- Verbatims vs scores: which answers which question
- How to read 500 comments in an afternoon
- Where text analytics helps and where it flattens
- The ten comments rule
- What quietly breaks the pairing
- Where to start
- FAQ
The satisfaction score for the support channel fell three points last quarter. The meeting spends forty minutes on it. Is it statistically significant? Did the sample change? Is it seasonal? Someone pulls up the previous four quarters. Someone else suggests waiting for another quarter of data before doing anything. Nobody in the room has read a comment, which is the oldest failure in handling qualitative and quantitative customer feedback.
Qualitative and quantitative customer feedback are the two halves of what customers tell you: the scores and counts that can be charted, and the comments and conversations that carry the reasons. The survey that produced the score also produced several hundred free-text answers, sitting in a column nobody opened.
Those comments are the only thing in the data set that can say what happened. The score cannot. A score that fell three points is a location, not a diagnosis. It tells you where to look. It does not tell you what you will find.
Key takeaways
- A score that moved is a location, not a diagnosis; only the comments behind it can say what happened.
- Quantitative feedback narrows the search to a channel, a quarter or a segment, and qualitative feedback finishes it.
- Five hundred verbatims can be read properly in an afternoon by sorting on score and reading the two extremes first.
- Text analytics counts consistently at scale and loses the detail you could act on, so let the tool count and let people read.
- No metric change should be presented without ten verbatims behind it, and failing to find ten is a finding in itself.
What qualitative and quantitative customer feedback are
The two words get used loosely, so here is the working distinction.
Quantitative feedback is the numbers: satisfaction and recommendation scores, retention rates, effort scores, contact volumes, resolution times, the count of complaints by category. It is comparable over time, it can be cut by segment, and it fits on a chart. Its weakness is that it has been stripped of everything that made it mean something. A seven out of ten does not say what the customer was thinking about when they chose it.
Qualitative feedback is the words: survey verbatims, interview notes, support transcripts, chat logs, reviews, the things frontline staff hear and occasionally write down. It is messy, uneven and hard to summarize. Its strength is that it still contains the reasons. A single comment can name the phone menu, the returns policy or the driver who left the parcel in the rain.
A survey with a score question and a free-text box produces both at once, tied to the same customer, which is the most useful arrangement there is and the one most often wasted. The comment gets counted as “one response with text” and the score gets averaged, and the link between them, the only thing that explains the average, is never used.
What neither is: a substitute for the other. Numbers without words produce meetings about significance. Words without numbers produce anecdotes, and the loudest anecdote wins.
Why scores alone lead meetings astray
The mistake is to treat the two as separate disciplines, one for the analysts and one for the researchers. They are two halves of the same question. The score narrows the search: this channel, this quarter, this segment. The comments finish it: the new phone menu, the policy change on returns, the two weeks when the outsourced team was short-staffed.
Working the other way round is just as useful. A theme that keeps appearing in comments is a hypothesis; the hard data tests whether it moves anything. A complaint that is loud but confined to a segment that is not churning can wait. One that is quiet but appears in the accounts you are losing cannot. The mildly worded comment from a long-standing customer is often the one that matters most, for reasons the piece on customers who leave without saying so explains.
This is also why customer analysis is harder to hand off than it looks. The question of whether the IT department can do your customer analysis usually comes down to this: the query is the easy part. Reading what the customers said and knowing what to make of it is the job.
Verbatims vs scores: which answers which question
Each half is good at some questions and useless at others, and most wasted effort comes from asking the wrong half.
| Question | Scores answer it | Verbatims answer it |
|---|---|---|
| Did something change? | Yes, if the sample held steady | Only if you read a lot of them |
| Where did it change? | Yes, by channel, segment or period | Roughly, and slowly |
| Why did it change? | No | Yes, usually within fifty comments |
| How many customers does it affect? | Yes, once you know what “it” is | No, comment counts are not prevalence |
| What exactly should we fix? | No | Yes, in the customer’s own words |
| Is it getting better after the fix? | Yes, next wave | Yes, and sooner, if the same theme fades |
One row deserves a note. Comment counts are not prevalence: a theme that appears in a hundred comments may reflect a hundred people or one issue that the survey question happened to prompt. The score across the affected segment is the only honest measure of size.
How to read 500 comments in an afternoon
People avoid reading verbatims because they imagine it takes weeks. Five hundred comments, read properly, takes an afternoon, and the method matters more than the stamina.
- Sort by score first. Put the comments next to the number the same customer gave. The comments from your lowest scorers and your highest scorers carry the most signal per sentence, so read the two extremes first: the bottom fifty and the top fifty. Within an hour you will have a feel for what is driving both ends. The top fifty will also include some of your best customers, who are otherwise the least heard people in the data.
- Tally themes by hand. Keep a sheet open, and every time a comment mentions something, make a mark under a theme. Themes should be specific enough to act on: “waiting for a callback” rather than “service”. You will start with five themes and end with fifteen. Merge the ones that turn out to be the same thing.
- Then read the middle. The middling scores are where the mixed feelings live, and where you find the customers who are one bad contact away from the bottom group.
- Copy the best sentences. Every theme should carry two or three comments quoted exactly, never paraphrased. A verbatim from a real customer does more in a management meeting than any bar chart, precisely because it has not been tidied up.
- Note the occasion. Where a comment makes it clear what kind of visit it was, a rushed reorder, a gift, a first purchase, mark that too. Reading comments occasion by occasion is how the same customer turns out to be several, and it explains many scores that look contradictory.
Only after all that should you open a text-analytics tool, if you have one. Now you know what it should find, and you will notice when it does not.
A worked example (illustrative)
Suppose the support survey produced 500 comments and the score fell from 78 to 75. Sorted by score, the bottom 50 comments mention “waited for a callback that never came” 22 times, “had to repeat myself” 14 times and a scattering of other things. The top 50 mention a named agent or “sorted it in one call” 30 times.
That is already a diagnosis: the callback queue broke, and first-contact resolution is what the happy customers are describing. Reading the middle 400 adds a third theme, “the new menu sent me to the wrong team”, which appears in mild comments at every score level and would never have shown up in the extremes. Total reading time is about three hours. The meeting that follows spends five minutes on significance and thirty on the callback queue.
Where text analytics helps and where it flattens
Text analytics is good at scale and consistency. Ten thousand comments a month cannot be read by hand, and a tool will tally themes the same way in March as in September, which a rotating team of readers will not. It is useful for tracking a theme over time once you know what the theme is, and for surfacing a sudden spike in a word or phrase you were not watching.
It flattens in three places. It collapses specific complaints into general categories, so “the driver left it in the rain” becomes “delivery: negative” and the detail you could act on is gone. It struggles with sarcasm, with comments that praise one thing and condemn another in the same sentence, and with the mildly worded message that is actually a customer giving up. And it produces a tidy output that looks finished, which discourages anyone from reading further.
The rule of thumb: let the tool count, and let people read. Sample the tool’s categories every month by pulling twenty comments from each and checking they belong. If a category turns out to be a bin of unrelated things, split it.
One more caution. Verbatims come with the same data-quality problems as any other customer record: duplicate responses, test entries, comments attached to the wrong contact. The argument that inaccurate data should not stop you from acting applies here too. Clean what you can, and read anyway.
The ten comments rule
A rule I have used with teams for years, and would happily have engraved somewhere: never present a metric change without ten verbatims behind it.
If the score fell, bring ten comments from the customers who pulled it down. If it rose, bring ten from the customers who lifted it. If you cannot find ten comments that explain the movement, say so in the meeting, because that is a finding in itself. It usually means the sample shifted rather than the experience.
The rule changes the meeting. Instead of forty minutes on significance testing, the discussion moves to whether the callback queue can be fixed by the end of the month.
The tally sheet that supports it can be as plain as this:
| Column | What goes in it |
|---|---|
| Theme | A specific, actionable label, such as “waiting for a callback” |
| Low-score mentions | Count of comments from the bottom of the scale that raise it |
| High-score mentions | Count from the top of the scale, because some themes appear at both ends |
| Two verbatims | Quoted exactly, with the score and segment noted |
| Owner | The team that could change it, or “nobody yet” |
Keep it in a shared file, update it with every survey wave, and it becomes the running diagnosis that the scores alone never were. The “nobody yet” column is worth watching on its own: a theme that sits there for three waves is a decision the company is declining to make.
What quietly breaks the pairing
The pairing of scores and comments fails in ways that look like diligence.
Comments get sampled to confirm. Someone reads twenty comments, finds the three that support the theory already in the room, and presents them. The fix is the sort-by-score method: the extremes first, all of them, before anyone forms a view.
Verbatims get paraphrased. “Customers feel the wait times are too long” is not a verbatim. It is an analyst’s sentence with a customer’s name on it, and it has lost the detail (how long, waiting for what, which channel) that made the original useful.
The two halves live in different teams. The analysts own the dashboard and the researchers own the interview notes, and they present in different meetings. Whoever presents the score should be the person who read the comments, or should at least sit next to them.
Comment counts get reported as prevalence. Twenty-two mentions of a callback problem is a strong signal, not a percentage of customers. The size of the problem comes from the score in the affected segment, not from the tally.
Anonymity strips the link. A survey that discards the connection between score and comment in the name of privacy has thrown away the one thing that made the data explain itself. Keep the link and protect the file; the two are not in conflict.
When comments are not enough
There are limits. A hundred responses with twenty comments cannot support the method, and the honest move is to say the wave is too thin and go and talk to ten customers instead. A theme found in comments is a hypothesis about cause, not proof; where a fix is expensive, the case for measuring it with a control group is the same as for any other intervention. And comments describe the customers who wrote them, who are more engaged and more articulate than the ones who did not, so a theme that is absent from the comments is not necessarily absent from the base.
Where to start
- Open the comment column from the last survey wave and put it next to the score column in one sheet.
- Read the bottom fifty and the top fifty this week, tallying themes by hand with specific labels.
- Copy two exact quotes per theme into a tally sheet with a “low”, “high” and “owner” column.
- Bring ten verbatims to the next metrics meeting, one page, behind whichever score moved most.
- Sample your text-analytics categories, if you have any, by pulling twenty comments from each and checking they belong.
- Write the ten comments rule into the reporting template so it applies to every score on every slide from now on.
FAQ
What is the difference between qualitative and quantitative customer feedback?
Quantitative feedback is the numbers customers give or generate: satisfaction and recommendation scores, effort ratings, contact volumes and retention rates. Qualitative feedback is the words: survey comments, interview notes, support transcripts and reviews. Numbers say where and how much; words say why and what exactly.
How do you analyze open-ended survey responses?
Sort the comments by the score the same customer gave and read the lowest and highest fifty first, tallying specific themes by hand. Then read the middle scores, copy two or three exact quotes per theme, and note who could act on each one. Five hundred comments read this way take about an afternoon.
How many customer comments do you need to find the reason for a score change?
The reasons usually show up within the first fifty comments at each end of the scale, because reasons cluster. Reading more improves the tally and turns up milder themes that hide in the middle scores. If ten comments cannot be found that explain a movement, the sample probably shifted rather than the experience.
Is text analytics better than reading comments by hand?
It is better at consistency and scale: a tool tallies the same way every month and can handle volumes nobody could read. It is worse at detail, sarcasm, mixed comments and the quiet message from a customer who is giving up. The practical rule is to let the tool count and let people read, checking the tool’s categories against a sample of twenty comments each month.
Are comment counts the same as how many customers are affected?
No. A theme mentioned in a hundred comments may reflect a hundred customers or one issue the question happened to prompt, and the people who write comments are more engaged than the people who do not. Use the score across the affected segment to size a problem, and use the comment count to prioritize what to read.