Testing Whether QA Scores Predict Anything
The analysis that validates or invalidates an entire quality programme, and it can be run in a week with data most operations already have.
A QA score claims to measure quality. Whether it does is testable, and the test is rarely run because the answer might be inconvenient.
The analysis
Join QA scores to outcomes at the interaction level. Same call: what was it scored, and what happened afterwards.
Outcomes to use: repeat contact within seven days, escalation, customer survey response, sale completed, complaint raised.
Look at whether they relate. Do higher-scored calls have better outcomes?
Do it per item, not only on the total. The total may be flat while individual items separate strongly, and those items are the ones worth keeping.
What operations typically find
The total score relates weakly to outcomes, more weakly than expected.
A small number of items carry most of the relationship. Frequently items about setting expectations, confirming understanding and diagnostic questioning.
Many items show no relationship at all, including some that are heavily weighted.
Tone and courtesy items relate to survey scores and weakly to repeat contact, which is a meaningful distinction: they affect how the customer felt, not whether the problem was solved.
Compliance items relate to nothing operational, correctly — they exist for regulatory reasons, not to improve outcomes, which is the argument for scoring them separately.
The confounders
Call complexity. Difficult calls score lower and have worse outcomes, so the relationship is partly driven by difficulty. Control by comparing within call category.
Agent tenure, which correlates with both.
Evaluator effect. If some evaluators score harder and are assigned to particular teams, the evaluator confounds everything. Control by including evaluator in the analysis, or by blinding.
Reverse causation on some items: a call going badly produces both a lower score and a worse outcome, without the scored behaviour causing anything.
The finding is an association, and it is far better than the alternative, which is no evidence at all.
Running it without a data team
Take four hundred evaluated calls from the last quarter.
Look up whether each produced a repeat contact within seven days.
Split into repeat and no-repeat.
Compare the average score on each item between the two groups.
Items where the groups differ noticeably are earning their place. Items where they do not are candidates for removal.
A spreadsheet and an afternoon. No statistical sophistication required to see a difference large enough to act on.
What to do with the result
Remove or demote items with no relationship, unless they exist for compliance reasons, in which case move them to the compliance section.
Increase the weight on items that separate the groups.
Rewrite items that are plausible but showed nothing — the item may be right and the definition wrong.
Re-run annually. What predicts outcomes changes.
The uncomfortable outcome
Sometimes the analysis finds that the QA score relates to nothing measurable.
That is a finding about the scorecard, not about quality. It means the form is measuring compliance with itself.
The response is to rebuild it from outcomes, which is described in its own note.
The response is not to stop measuring, and it is not to bury the analysis — though that is the common outcome, because the programme has been reporting the score for years.
Handling a null result
Sometimes the analysis finds that the scorecard predicts nothing, and what happens next determines the programme's integrity.
Check the analysis first. Insufficient sample, outcome measured wrongly, evaluator effect swamping the signal, or all calls in one category.
Check the outcome measure. Repeat contact with a poor matching rule is noisy enough to hide a real relationship.
Check the score distribution. If it is compressed, there is not enough variation for anything to correlate with.
If the analysis holds, report it. The scorecard measures compliance with itself.
Rebuild from outcomes rather than adjusting weights on items that predict nothing.
Do not bury it. The finding will be rediscovered, and a programme that suppressed it has a much larger problem than a scorecard that needed rebuilding.