Why Two Percent Tells You Nothing
The standard practice of scoring a handful of calls per agent per month produces a number with enormous uncertainty, presented to two decimal places.
A typical QA programme scores between two and six interactions per agent per month. The resulting percentage is treated as a measurement of that agent's quality. The statistics do not support that.
The arithmetic
An agent handles perhaps four hundred calls a month. You score four.
That is a one percent sample, and it is used to produce a score compared against a target, reported to management, and in many organisations linked to pay.
The confidence interval on a proportion estimated from four observations is enormous. An agent scoring three out of four on some behaviour has a true rate that could plausibly be anywhere from well under half to nearly all.
Month-to-month variation in such a score is mostly noise. An agent moving from 88 to 94 has probably not changed.
What this produces in practice
Coaching aimed at randomness. A supervisor discusses the specific failings in four calls, which may be unrepresentative.
Rankings that reshuffle monthly and are read as performance movement.
Regression to the mean read as improvement. Coach the lowest scorers and they rise next month, because part of their low score was noise. This makes every intervention look effective.
Disputes that are statistically justified. An agent arguing that the four calls were unrepresentative is frequently correct.
What would be needed
To detect a moderate difference between agents with reasonable confidence requires far more observations than any manual programme produces — the number depends on the behaviour's base rate, and for anything that occurs on most calls it runs into dozens per agent per period.
For rare behaviours it is worse. A compliance failure occurring on two percent of calls cannot be estimated from four calls at all.
Manual QA cannot get there. Evaluator time is the constraint, and doubling the sample doubles the cost for a modest reduction in uncertainty.
What to do instead
Stop treating the agent score as a measurement. Treat it as a set of observations for coaching, which is what it is good for. Four calls give you four concrete examples to discuss, and that has value independent of any statistic.
Aggregate across agents for anything you intend to measure. A team's or a site's score from two hundred observations is a far more defensible number than an individual's from four.
Use analytics for the things that need coverage. Compliance presence checks, category counts, acoustic measures — these can run on every call and the sampling problem disappears.
Reserve human evaluation for judgement, and select the calls deliberately rather than randomly: calls flagged by analytics, calls with poor outcomes, calls from a specific process you are investigating.
Report uncertainty. If an agent score is reported, report the number of observations next to it. A percentage from four calls displayed as "91.7%" is a false precision that shapes decisions.
The pay and performance question
Where QA scores feed compensation or performance management, the sampling problem becomes an employment issue.
A score with wide uncertainty, used to differentiate between agents, will produce unfair outcomes and they will be challenged.
The defensible arrangements: use QA for development rather than for ranking; where it must feed performance, aggregate over longer periods and larger samples; and always allow the agent to see the specific calls and dispute them.
A programme that cannot show which calls produced a score cannot defend the score, and that is a common state of affairs.
Reporting a score with its sample size
A small change in presentation that prevents most misreading.
Always show the number of evaluations next to any agent-level score. "91% (4 evaluations)" reads very differently from "91%".
Show the range across those evaluations, not just the mean. Four calls scoring 70, 95, 98 and 100 average to 91 and describe something other than consistency.
Round to the nearest five at small samples. False precision invites the belief that a two-point difference means something.
Report team and site scores with their sample sizes too, which are large enough to be meaningful and are frequently buried under the individual figures.
Where a score feeds any consequence, state the sample size in the same sentence. An agent challenged on a percentage derived from four calls should be able to see that it was four.