What Contact Centre QA Actually Is
Three activities share the name: verifying compliance, improving performance, and understanding customers. They need different scorecards.
Ask three people in a contact centre what quality assurance is for and you get three answers. All are legitimate and they pull in different directions, which is why so many QA programmes produce numbers nobody acts on.
The three purposes
Compliance verification. Did the agent say the required disclosure, verify identity correctly, follow the regulated script? Binary, auditable, and the reason QA exists in regulated industries.
Performance improvement. Is the agent doing the job well, and where do they need help? Judgement-based, developmental, and the reason agents care.
Customer and process insight. What are customers actually calling about, what is going wrong upstream, what does the business need to fix? Aggregate, analytical, and the highest-value output of the three.
A single scorecard cannot serve all three, and most organisations try, which produces a form with forty items where the compliance checks and the coaching observations sit in the same weighting.
What a QA programme physically consists of
A sample of interactions — calls, chats, emails, tickets.
A scorecard defining what is assessed.
Evaluators applying it.
A score, usually a percentage.
A feedback mechanism back to the agent.
Reporting to whoever asked for it.
Each of those six is a design decision and each is usually inherited rather than chosen.
Where speech analytics fits
Speech analytics does not replace this. It changes the sampling problem.
Manual QA reviews a few interactions per agent per month — typically two to six. That is a vanishingly small sample of an agent's work, and the statistics of it are covered separately because they are worse than people assume.
Speech analytics processes everything. It cannot judge nuance, empathy or whether the resolution was correct. It can find every call where a required phrase was absent, where the customer became distressed, where a competitor was mentioned, or where hold time exceeded a threshold.
The combination is the point: analytics for coverage and detection, human evaluation for judgement. Programmes that treat analytics as automated QA produce a compliance system with an opinion about tone.
The failure that defines most programmes
QA becomes a score, and the score becomes a target, and the target becomes the point.
Agents optimise for the scorecard. Evaluators drift toward scores that avoid arguments. The number stabilises around a comfortable value and stops carrying information.
The tell is a score distribution with almost no variance. If every agent scores between 92 and 97, the instrument is not measuring anything. This is extremely common and it is treated as evidence that quality is high.
What a working programme looks like
Separate scorecards for separate purposes, or at least separate sections with separate reporting.
Compliance items scored as pass or fail, not weighted into a percentage where failing one can be offset by tone.
A small number of behavioural items that someone has evidence actually relate to outcomes.
Calibration between evaluators, run regularly, with the disagreement measured.
Analytics providing coverage so that the human sample can be chosen rather than random.
Findings routed to process owners, not only to agents. A large share of what QA discovers is not an agent problem.
The question worth asking first
What decision does this score inform?
If the answer is "we report it monthly", the programme is producing a number rather than an outcome. If the answer names a coaching action, a process fix or a compliance obligation, the scorecard can be designed backwards from that, which is the only way it ends up measuring something useful.
The first diagnostic on an existing programme
Five questions that establish whether a QA programme is producing information or a number.
What is the score distribution? If almost everyone sits within a few points, the instrument has stopped discriminating.
What is the inter-rater agreement? If it has never been measured, no score in the programme has a known reliability.
How many interactions per agent per period? If it is a handful, the individual score has a confidence interval nobody has calculated.
When was the scorecard last changed, and why? A form untouched for three years is measuring a theory of quality nobody has tested.
Name three things the programme caused to be fixed. If nobody can, the output is going to agents and nowhere else, which is the most common failure mode.
A programme that answers all five well is unusual. One that cannot answer any of them is producing a monthly percentage that has been reported for years without anyone checking whether it means anything.