Survey Scores and What They Actually Measure
Response rates are low and non-random, so a survey average describes the people who answered. What that supports and what it does not.
Customer surveys are the closest thing to a direct quality measure and they carry a selection problem that most reporting ignores.
The response rate problem
Typical post-contact response rates are low, frequently in the single digits to low tens of percent.
Responders are not a random sample. People with strong opinions in either direction respond disproportionately; the indifferent majority does not.
This produces a bimodal distribution and an average that sits between two peaks where few actual responses fall.
Reporting the average conceals the shape. A score of 4.1 from a distribution of mostly fives and some ones describes a different situation from a genuine cluster around four.
Report the distribution, or at minimum the proportion of top-box and bottom-box responses.
What the surveys measure
Post-contact satisfaction measures the interaction, roughly, and is contaminated by the outcome. A customer told no rates the agent lower regardless of how well it was delivered.
Relationship measures ask about the company, not the contact, and using them to evaluate an agent is a category error that happens constantly.
Effort measures — how easy was it — correlate better with repeat contact and loyalty than satisfaction does, according to the widely cited research, and are less contaminated by the outcome.
Recommendation measures are about the brand. Attributing them to an individual interaction is unsupportable.
The attribution problem
A survey response reflects the outcome, the product, the price, the wait, the policy and the agent.
Attributing it to the agent is the standard practice and it is wrong in proportion to how much the other factors dominate.
Agents on queues that deliver bad news score lower, systematically and permanently, and it is not about them.
Normalise by queue and by outcome before comparing agents, or do not compare agents on it at all.
Using it with QA data
The pairing is where the value is.
Calls with low survey scores are a review queue. Better than random selection and better than analytics flags alone.
Calls scoring well on QA and badly on survey are the interesting ones. The agent followed the scorecard and the customer was unhappy, which usually means the scorecard is measuring the wrong thing or the outcome was the problem.
Calls scoring badly on QA and well on survey are equally interesting and usually mean the agent did something effective that the scorecard does not credit.
Both categories are a source of scorecard improvement and almost nobody looks at them.
Free text
The most valuable and least used part of the survey.
Free text responses are analysable with the same categorisation tools applied to calls, and the volume is small enough to read.
They contain the reason, which the score does not.
They frequently identify process problems more directly than call analysis, because the customer says what went wrong in their own words.
Categorise them and trend the categories alongside the call categories. Where the two disagree, one of them is missing something.
What not to do
Do not target survey scores at agent level without normalising for queue and outcome. It produces agents avoiding difficult calls and gaming for responses.
Do not compare across organisations. Methodology, timing, channel and question wording differ.
Do not report an average without the response rate and the distribution.
Do not read a small monthly change as a signal. With low response rates, the confidence interval on a team's monthly score is wide, and the same statistical argument applies as to QA sampling.
Reading the free text properly
The most underused survey output and the easiest to analyse.
Categorise it with the same tooling used on calls. Volume is low enough that precision can be high.
Compare the free-text categories to the call categories. Where they disagree, one channel is missing something.
Separate comments about the agent from comments about the outcome, the product and the wait. Most operations conflate them and attribute all of it to the agent.
Read a sample by hand every month. Nothing substitutes for it, and supervisors who do it describe it as the most useful half hour they spend.
Route product and policy comments to those owners, with counts.
Track the proportion mentioning something outside the agent's control. A high and rising figure means survey scores are measuring the operation rather than the people being scored on them.