Programme
Designing the programme
Scorecards, calibration, automation and rollout. Most of the avoidable failure in this field happens here, before any evaluation takes place.
Procedure
Defining Quality Before Measuring It
Most scorecards encode assumptions about what good looks like that nobody has tested. Deriving them from outcomes takes a quarter and changes what gets measured.
Procedure
Calibration: Keeping Evaluators Consistent
Two evaluators scoring the same call disagree more than anyone expects. Measuring the disagreement is the only way to know whether your scores mean anything.
Analysis
Automated Scoring: Where It Works
Scoring every interaction is the headline promise. It works for a narrow band of items and produces disputes everywhere else.
Analysis
What Changes When You Analyse Every Interaction
Coverage removes the sampling problem and introduces a volume problem. What genuinely improves, and what organisations discover they were not ready for.
Analysis
Targets, Gaming and Goodhart's Law
Any QA measure that becomes a target stops measuring what it did. The mechanisms are predictable and some are avoidable.
Procedure
Rolling Out a QA or Analytics Programme
The technical deployment is the easy part. Agent trust and the first month's findings determine whether the programme survives a year.
Analysis
Multi-Language and Multi-Site Programmes
Analytics quality varies enormously by language, which means a global programme measures some sites better than others and reports the difference as performance.
Analysis
Chat, Email and Messaging Quality
Text channels remove the transcription problem entirely and introduce their own. Most programmes apply a voice scorecard to them, which measures the wrong things.
8 notes in this section