Targets, Gaming and Goodhart's Law
Any QA measure that becomes a target stops measuring what it did. The mechanisms are predictable and some are avoidable.
A quality score used as a target is optimised toward, by agents, by evaluators and by managers. The optimisation is rational and it destroys the measure.
How each party games it
Agents learn the scorecard and perform to it. Saying the required phrase without meaning it. Front-loading the checklist. Avoiding the behaviours that get marked down rather than pursuing the outcomes.
Evaluators drift toward scores that avoid arguments. Scoring generously is faster and less unpleasant than defending a low mark.
Supervisors, whose teams' scores reflect on them, apply pressure in one direction.
Managers select which calls are reviewed where they have discretion.
None of this requires bad intent. Every participant is responding sensibly to how they are measured.
The signatures
Compressed distribution. Everyone between 90 and 97. The instrument has stopped discriminating.
Scores rising while outcomes do not. The clearest evidence of gaming and the easiest to check, and almost nobody checks.
Item-level saturation. An item nobody ever fails.
Perfect compliance with a phrase, no change in the outcome the phrase was meant to produce. The disclosure is said and customers still do not understand.
Disputes falling to zero, which usually means agents have stopped bothering rather than that scoring improved.
Reducing it
Never set a target on a measure you also use to understand the operation. Use different measures for management and for insight, or the insight measure becomes a management measure.
Anchor on outcomes. Targets on resolution, repeat contact and customer feedback are harder to game than targets on process compliance, because gaming them requires actually doing the job.
Measure the outcome alongside the score. A rising score with flat outcomes is the alarm.
Rotate what is emphasised. Optimising toward a stable target is easier than toward a shifting focus.
Random selection for review where possible, so calls cannot be chosen.
Blind evaluation to the agent's identity.
Full coverage, which removes selection entirely for automated items.
The compliance exception
Compliance items should be targeted, at 100 percent, and gaming them is called doing the job.
Saying the disclosure because it is scored is the intended outcome. The requirement is that it was said.
The failure mode is different: saying it so fast the customer cannot follow. This is worth checking specifically, and it is measurable — disclosure delivered in under a threshold duration.
The pay question, again
Linking a gameable measure to pay maximises the gaming.
The combination of a small sample, evaluator disagreement, and a financial consequence produces exactly the behaviour anyone would predict, and the disputes that follow are justified.
If QA must feed pay, use aggregate measures over long periods, include outcome measures, and keep the dispute route open and used.
Better: pay on outcomes, coach on quality. Then the QA conversation can be about improvement rather than about money, which is the only condition under which coaching works.
The check worth running annually
Plot QA scores against a hard outcome over two years.
If scores rose and outcomes did not, the measure has decayed and the scorecard needs rebuilding.
If they moved together, the measure is still carrying information.
This single chart tells you more about the health of a QA programme than any amount of process documentation, and almost no operation produces it.
The annual decay check
One chart, produced once a year, that tells you whether the QA measure still carries information.
Plot the average QA score by month over two years.
Plot an outcome measure on the same axis — repeat contact rate, inverted so that better is up.
Look at whether they move together.
Together: the score is still measuring something.
Score rising, outcome flat or worse: the measure has been optimised and no longer indicates quality. Rebuild the scorecard.
Both flat: either nothing is changing, or neither measure is sensitive enough. Check the score distribution for compression.
Add the score distribution width as a third series. Narrowing distribution with rising average is the clearest possible signature of gaming.
This chart takes an hour and almost no operation produces it, which is why decayed scorecards persist for years.