Customer Analytics

Customer experience measurement: is your program worth the effort?

Customer experience measurement decides whether a program survives year two. Rule 7 of no-excuses CX: three questions, a comparison group, one page a month.

Table of contents
  1. Key takeaways
  2. What customer experience measurement is (and what it is not)
  3. Why rule 7 protects the other six
  4. How to design customer experience measurement before launch
  5. Comparison group vs last year vs staggered start: which to use when
  6. Leading indicators inside the quarter
  7. What quietly breaks customer experience measurement
  8. When measurement is not the answer
  9. Where to start
  10. FAQ

This is part 4 of a five-part series on no-excuses customer experience, and the part about customer experience measurement. Part 1 introduced the excuses and the seven rules, part 2 covered rules 3 to 5 and part 3 covered rule 6. This part takes rule 7, and part 5 closes the series with lifetime value.

The program has been running for six months. The callbacks happened, the return path got two steps shorter, the frontline has its list. Then the CFO asks the question you knew was coming: was it worth it?

You have numbers. Calls made. Emails opened. A retention rate that is up a little on last year. But last year also had a price rise, a competitor’s outage and a new product line, and the CFO knows it. What you do not have is an answer, because the answer needed to be designed before the program started, and it was not. Customer experience measurement is the practice of showing, against a comparison, what a customer program changed in customer behavior and in money.

This is the most common way a good customer program ends: not disproved, just unproven, which in a budget meeting comes to the same thing.

Key takeaways

  • Rule 7, measure it or it did not happen, is what lets a customer program survive into its second year, because an unproven program is cut the first time money gets tight.
  • Measurement is designed before launch, while there is still a baseline and a chance to hold a comparison group aside.
  • Three questions settle the design: what would change, for whom, and compared with what.
  • A comparison group is the difference between evidence and anecdote; last year is a weak comparison because last year was different in a hundred ways.
  • Leading indicators inside the quarter keep a program alive while the real outcome takes shape, as long as nobody mistakes them for the answer.

What customer experience measurement is (and what it is not)

Measurement in this sense is not the same as reporting. Reporting counts what you did. Measurement shows what changed because of it, for a defined group of customers, relative to a group you did not touch. The three kinds of number a customer program produces are easy to confuse, and only one of them answers the CFO.

Kind of metric What it counts Example What it proves
Activity What the team did Calls made, emails sent, steps removed That the work happened
Perception What customers say Satisfaction score, survey comments How customers felt at the moment of asking
Outcome What customers did Still active a year later, repeat purchase, renewal completed Whether the program changed behavior, if there is a comparison

Activity metrics are fine as a check that the plan is being executed. Perception metrics are useful as an early sign and as a source of verbatims. Outcomes are the only kind that survive a budget meeting, and only when they sit next to a comparison group. Without the comparison, an outcome is just a trend with a story attached.

Why rule 7 protects the other six

Rule 7 is the last of the seven and it protects all the others. Rules 1 to 6 get a program started. Rule 7 is what lets it survive into its second year.

  1. Start with the customers you already have.
  2. Act on what you already know.
  3. Talk to customers like a person: their channel, their timing, their words.
  4. Make coming back easy: remove friction at the moment of return.
  5. Give the frontline a reason and a way.
  6. Sell it inside before you sell it outside.
  7. Measure it, or it did not happen.

“It did not happen” sounds harsh. It is meant to. A program that cannot show what it changed will be cut the first time money gets tight, and it will deserve to be, because nobody can tell it apart from a program that did nothing. The customers it helped will not come to the meeting to defend it.

The excuse here is “we cannot measure it”, and it is the most respectable of the six excuses because it is sometimes half true. You cannot measure everything. You can measure something, and the craft is in choosing what before you launch, while you still have a baseline and a chance to hold a comparison group aside.

How to design customer experience measurement before launch

Every measurement plan I trust answers the same three questions, and answers them in writing before the first customer is touched.

What would change?

Name the outcome, not the activity. Calls made is an activity. Customers still active twelve months after a closed ticket is an outcome. If the honest answer is “we hope retention improves”, push until it becomes “we expect repeat purchase in the treated group to be higher than in the untreated group within two quarters”.

For whom?

The population, defined by a rule you could apply with a spreadsheet filter. Not “at-risk customers” but “customers with a support ticket closed without a resolution note in the last thirty days”. If you cannot list them, you cannot measure them.

Compared with what?

This is the question that separates evidence from anecdote. Compared with last year is weak, because last year was different in a hundred ways. Compared with the customers you did not treat, chosen the same way, at the same time, is strong. That is a comparison group, and it is less complicated than its reputation.

The one-page plan, and the one-page monthly report

The final discipline is the report. One page. Monthly. Written so that a CFO could read it without you standing next to it. Outcome at the top, comparison group beside it, the leading indicators underneath, and one line of interpretation. No activity metrics unless they explain a movement in something that matters.

Here is the template. Fill it in before launch, and the page a month mostly writes itself.

Element What goes here Example
Outcome The thing that should change, in customer terms Still active twelve months after a closed ticket
Leading indicator A sign inside the quarter that the outcome is coming Purchase within thirty days of the callback
Baseline Where the number stood before you started Measured on the previous six months of comparable customers
Comparison Who you will compare against, and how they were chosen One in ten eligible customers, held out at random
Date to decide The day someone says scale, stop or change, and who says it End of the second quarter, the sponsor

The last row is the one most plans leave out, and it is the one that makes the rest matter. A measurement without a decision date is a report. A measurement with one is a program that knows why it exists.

Comparison group vs last year vs staggered start: which to use when

The objection to holding customers out is always the same: it feels wrong to withhold something good from a customer on purpose. I understand that, and I still argue for it, for two reasons.

First, you do not yet know it is good. That is what the pilot is for. Withholding an unproven intervention from a random tenth of eligible customers for a quarter is a smaller harm than running an unproven program on everyone for three years.

Second, without the comparison, the program has no defense. The next time someone says “retention was going up anyway”, you will have nothing to say. With the comparison, you have the one sentence that ends the argument: the customers we called stayed at a higher rate than the customers we did not, and the only difference was the call.

Method How it works What it proves Use it when
Random hold-out One in ten eligible customers, chosen at random, is not treated The program caused the difference Always, if you possibly can
Staggered start One region or segment is treated a quarter before the rest Probably the program, unless something else happened to that region Holding customers out is genuinely impossible
Last year This year’s number against the same period last year That something changed, cause unknown Only as context, never as the answer

If you truly cannot hold anyone out, use the staggered start: treat one region or one segment first and the rest a quarter later, and compare the gap. It is weaker, and it is much better than nothing.

Leading indicators inside the quarter

Lifetime outcomes take a long time to arrive, and budgets do not wait. So every measurement plan needs a second layer: indicators that move within the quarter and sit upstream of the outcome you care about.

For a callback program, that might be the share of called customers who make a purchase in the following thirty days. For a friction fix at renewal, the completion rate of the renewal flow and the volume of renewal calls to the service desk. For a frontline program, whether the tool is actually opened.

These are not the answer, and it matters to say so out loud. They are the early signs that let you keep going while the real answer takes shape, and they are what lets you work with lifetime value in a business that thinks in quarters.

A worked example (illustrative)

The numbers are round and invented, to show what the page says and what it does not.

A callback program treats 1,800 customers in a quarter and holds out 200 at random. The leading indicator is purchase within thirty days of the ticket closing. At the end of the quarter, 540 of the treated customers have bought (thirty in a hundred) against 48 of the held-out (twenty-four in a hundred). The one line of interpretation reads: “Called customers bought within thirty days at a higher rate than held-out customers; the twelve-month outcome is not yet available.” That is all the page claims.

What it does not say is that the program made money, because the outcome has not arrived and the margin on those purchases has not been counted. The temptation to write “revenue up” is strong and it is the fastest way to lose the CFO’s trust, because the next quarter will ask for the number that supports it. The leading indicator buys the program a second quarter. The outcome, and the value of it, is what buys the second year.

What quietly breaks customer experience measurement

The comparison group is contaminated. A well-meaning agent sees a held-out customer on a list and calls them anyway. After a quarter of that, the two groups are the same group. The hold-out list needs to be as protected as the treatment list, and the frontline needs to know why.

The population rule changes mid-way. “Tickets closed without resolution” becomes “tickets closed without resolution, plus anyone who complained on social media”. The numerator and denominator no longer match the baseline, and the comparison is gone. Change the rule for the next cohort, not the current one.

Retention is counted the flattering way. If “still active” means “still in the database” rather than “bought something”, every group looks retained. Measuring retention correctly is a precondition for measuring the program.

The measurement becomes a report request. Handed to IT as a ticket, the plan comes back as a dashboard of activity metrics, because that is what the systems hold. Someone who understands the customer question has to own the design, which is the argument for not leaving customer analysis to the IT department.

When measurement is not the answer

Rule 7 is strict, and there are cases where it should bend.

When the population is too small to compare. A program that touches forty customers a quarter will not produce a comparison anyone can trust. Measure the process instead (did the action happen, within how many days) and collect verbatims, and be honest that the outcome will be judged qualitatively.

When you would do it anyway. Fixing a broken renewal page, or calling a customer whose complaint was lost, is sometimes simply the right thing to do. Do it, note that it was not measured, and spend the measurement effort on the decisions that are genuinely open.

When the outcome is years away. For long-tenure relationships the twelve-month outcome may itself be a leading indicator. Say so on the page, and set the decision date on the best proxy you have rather than waiting for a number that will arrive after the budget cycle has moved on.

Where to start

  1. Write the three answers on one page before launch. What would change, for whom, compared with what. If the program has already started, write them now for the next cohort.
  2. Define the population as a filter rule. Something a colleague could apply to the customer list without asking you.
  3. Hold out one in ten at random and protect the list. Tell the frontline what the list is for.
  4. Pick one leading indicator that moves within thirty days and sits upstream of the outcome.
  5. Book the decision meeting now. A date, a room, the sponsor and the people who would fund scaling, with the rule for scale, stop or change written on the invitation.

FAQ

What is customer experience measurement?

Customer experience measurement is the practice of showing what a customer program changed in customer behavior and money, for a defined group of customers, relative to a comparable group that was not treated. It is different from reporting, which counts activity such as calls made or emails sent. The core of it is three questions answered before launch: what would change, for whom, and compared with what.

How do you measure whether a customer experience program worked?

Define the outcome in customer terms, such as still buying twelve months later, and the population by a rule you can apply as a filter. Hold out a random share of eligible customers and do not treat them. Compare the outcome between the treated and held-out groups on a decision date that was set before launch.

Do you need a control group to measure customer experience?

A random hold-out group is the strongest evidence and the only one that answers “was it going up anyway”. Where holding customers out is impossible, a staggered start (treating one region or segment a quarter before the rest) is a weaker but usable substitute. Comparing with last year shows that something changed but not what caused it.

What is a leading indicator in customer experience measurement?

A leading indicator is a number that moves within the quarter and sits upstream of the outcome you care about, such as the share of called customers who buy within thirty days. It lets a program show early signs of working while the real outcome, such as twelve-month retention, takes shape. It is an early sign rather than the answer, and the monthly page should say so.

How often should you report on a customer program?

Once a month, on one page, written so that a finance leader could read it without the author present. The page shows the outcome, the comparison group, the leading indicators and one line of interpretation. Activity metrics belong on it only when they explain a movement in something that matters.

Related articles