Quality · Analytics · Monitoring
The transcript is not the conversation
Everything speech analytics produces sits on a transcript that is imperfect, and imperfect unevenly. These fifty notes cover what the tools measure, what they cannot, how to build a quality programme that survives contact with agents, and where the monitoring obligations bite.
What this is
Fifty notes on contact centre quality assurance and speech analytics, written for the people who run these programmes rather than for the people who sell the platforms.
Independent notes and software comparisons, with no sponsored placements, affiliate arrangements or tracking scripts.
Four things that hold across every deployment
Transcription accuracy sets the ceiling. Every category, every automated score and every sentiment figure rests on text that is imperfect. Measure the error rate on your own calls before building anything on it.
It is imperfect unevenly. Speech recognition performs measurably worse for some speakers than others, along documented lines. Where automated scoring reaches an agent's record, that is an employment matter and it needs measuring.
Small samples do not measure people. Four calls a month, scored by evaluators who have never had their agreement measured, is not a characterisation of an agent's work.
A large share of findings are not about agents. Long silences, repeated questions and clustered compliance failures are systems, knowledge and policy problems. Routing them correctly is the highest-value output of the whole exercise.
Where to start
Setting up a programme: the foundations, then defining quality from outcomes, then calibration.
Already running one: the sampling note and the test of whether your scores predict anything.
About to buy analytics: measure your transcription accuracy first, and check whether your recordings are stereo.
Foundations
Three different activities share the name quality assurance, and they need different instruments. Plus the sampling arithmetic that undermines most agent-level scoring.
Explainer
What Contact Centre QA Actually Is
Three activities share the name: verifying compliance, improving performance, and understanding customers. They need different scorecards.
Explainer
What Speech Analytics Actually Does
Transcription, then search and classification over the transcript. Everything vendors describe as understanding is pattern matching on text that is itself imperfect.
Analysis
Why Two Percent Tells You Nothing
The standard practice of scoring a handful of calls per agent per month produces a number with enormous uncertainty, presented to two decimal places.
Procedure
Anatomy of a Scorecard That Measures Something
Most scorecards accumulate items nobody removes, weight them arbitrarily, and produce a percentage that cannot be traced to any outcome. A shorter one works better.
Analysis
Manual and Automated Evaluation: What Each Catches
They fail in opposite directions. Automation misses meaning and covers everything; humans understand meaning and see almost nothing.
Analysis
QA, Coaching and Compliance Are Different Functions
Combining them into one process is the most common structural mistake in contact centre quality, and it makes the coaching adversarial and the compliance unreliable.
Reference
What This Data Cannot Tell You
A list of questions routinely asked of quality and speech analytics data that it cannot answer, with what to use for each.
How the analytics works
Transcription accuracy sets the ceiling for everything downstream, and it varies systematically by speaker. This section covers what each component actually does and where it fails.
Reference
Transcription Accuracy and What Degrades It
Vendor accuracy figures come from clean benchmark audio. Contact centre calls are not that, and the difference is large enough to change what the system can support.
Analysis
The Accent and Dialect Accuracy Gap
Speech recognition performs measurably worse for some speakers, along lines that map onto race and region. That has consequences.
Procedure
Categorisation and Topic Detection
The most useful thing speech analytics does, and the thing most often set up badly. Category definitions are the whole product and they need maintenance.
Analysis
Sentiment Analysis: What It Measures
It classifies language as positive or negative. That is not the same as knowing how the customer felt, and the difference matters.
Analysis
Emotion Detection and the Science Problem
Products claim to infer emotional state from voice. The scientific foundation is contested and regulators have begun restricting it.
Reference
Acoustic Measures: Silence, Talk-Over and Pace
The most reliable outputs in the whole toolkit, because they are measured rather than inferred, and the least discussed.
Procedure
Redaction, PCI and Sensitive Data in Recordings
Recordings capture card numbers, health data and identifiers. Redaction is a control with known failure modes, and so is pause-and-resume.
Analysis
Real-Time Analytics and Agent Assist
Prompting during a live call is a different product with different failure modes. When it helps, when it distracts, and the surveillance question it raises.
Analysis
Language Models in Conversation Analysis
Vendors have replaced rule engines with language models. This genuinely improves some tasks, introduces new failure modes, and makes the system harder to audit.
Reference
Keyword Spotting and Full Transcription
Two architectures with different economics and a decisive practical difference: whether you can ask a new question of last year's calls.
Reference
Speaker Separation and Why Stereo Matters
Knowing who said what is the foundation of every agent-level measure. Mono recording makes it a guess, and a large number of deployments discover this late.
Programme design
Scorecards, calibration, automation and rollout. Most of the avoidable failure in this field happens here, before any evaluation takes place.
Procedure
Defining Quality Before Measuring It
Most scorecards encode assumptions about what good looks like that nobody has tested. Deriving them from outcomes takes a quarter and changes what gets measured.
Procedure
Calibration: Keeping Evaluators Consistent
Two evaluators scoring the same call disagree more than anyone expects. Measuring the disagreement is the only way to know whether your scores mean anything.
Analysis
Automated Scoring: Where It Works
Scoring every interaction is the headline promise. It works for a narrow band of items and produces disputes everywhere else.
Analysis
What Changes When You Analyse Every Interaction
Coverage removes the sampling problem and introduces a volume problem. What genuinely improves, and what organisations discover they were not ready for.
Analysis
Targets, Gaming and Goodhart's Law
Any QA measure that becomes a target stops measuring what it did. The mechanisms are predictable and some are avoidable.
Procedure
Rolling Out a QA or Analytics Programme
The technical deployment is the easy part. Agent trust and the first month's findings determine whether the programme survives a year.
Analysis
Multi-Language and Multi-Site Programmes
Analytics quality varies enormously by language, which means a global programme measures some sites better than others and reports the difference as performance.
Analysis
Chat, Email and Messaging Quality
Text channels remove the transcription problem entirely and introduce their own. Most programmes apply a voice scorecard to them, which measures the wrong things.
Running it
Feedback, disputes, evaluator workload and the routing of process findings — which is where the largest value sits and where most programmes stop.
Procedure
Delivering Feedback That Changes Behaviour
The evaluation is worthless if the conversation produces no change. What the research suggests, and what practice usually does instead.
Procedure
A dispute route is not an administrative burden. It is the mechanism that finds broken scorecard items, inconsistent evaluators and transcription failures.
Procedure
Getting Process Findings Out of QA Data
Much of what QA discovers is not an agent problem. Routing those findings to people who can fix them is the highest-value output.
Analysis
Evaluator Workload and Score Quality
Score quality degrades measurably with volume and session length. Most programmes set evaluator targets without knowing where that threshold sits.
Analysis
QA in Outsourced and BPO Environments
When the agents work for someone else, QA becomes a contractual instrument as well as a quality one, and the incentives point in unhelpful directions.
Procedure
From selecting an interaction to a closed action. Most programmes have the first three steps and stop, which is why the output does not change anything.
Analysis
What QA Does to the People Being Measured
Contact centre attrition is high and quality programmes contribute to it. The design choices that make monitoring tolerable are known and rarely applied.
Measurement
The short set of measures that carry information, the test that validates a whole scorecard, and how to report to people who want one number.
Reference
A short set that carries information, and the longer set that fills dashboards and drives the wrong behaviour.
Procedure
Testing Whether QA Scores Predict Anything
The analysis that validates or invalidates an entire quality programme, and it can be run in a week with data most operations already have.
Analysis
First Contact Resolution, Honestly
The most quoted contact centre metric and the most inconsistently defined. Most published FCR figures are not measuring resolution.
Analysis
Survey Scores and What They Actually Measure
Response rates are low and non-random, so a survey average describes the people who answered. What that supports and what it does not.
Procedure
Detecting Scorecard and Model Drift
Scores move for reasons unrelated to performance: evaluator turnover, model updates, call mix. Instrumenting distinguishes them.
Procedure
Reporting Quality to Leadership
A single quality percentage is what gets asked for and it is the least informative thing you can supply. What to report instead, and how to survive the request.
Compliance and ethics
Recording consent, biometric templates, employee monitoring obligations, retention, and the measured accuracy disparities that make automated scoring an employment matter.
Reference
Recording rules differ by jurisdiction, by party and by purpose. A single global recording policy will be wrong somewhere, and the exposure is real.
Reference
Voice authentication creates a biometric identifier, which is regulated more strictly than a recording. Several statutes carry private rights of action.
Reference
Monitoring Agents: The Employment Dimension
Continuous analysis of employee speech is employee monitoring, which in many jurisdictions carries notification, consultation and proportionality obligations.
Procedure
Retention: Recordings, Transcripts and Indexes
Deleting a recording deletes one copy. The transcript, the analytics index, the summary and the backup are separate objects and they are usually forgotten.
Analysis
Bias in Automated Quality Scoring
Automated scoring inherits the speech recognition accuracy gap, so it can be systematically harsher toward some agents. Measure it.
Checklist
Agents will find out what the system does. Telling them first is a legal requirement in several jurisdictions and an operational advantage everywhere.
Checklist
Vendor Due Diligence on Your Recordings
Your recordings leave your estate. Where they go, who hears them, how long they stay and whether they train models are contract questions.
Buying and starting
What to test on your own audio, what to ask about your data, and a first quarter that costs nothing and produces a programme worth automating.
Checklist
Evaluating a Speech Analytics Platform
Demonstrations run on clean audio with tuned categories. The questions and tests that reveal what the product does on your calls.
Analysis
Building Speech Analytics Rather Than Buying
Open transcription models have made the technical build tractable. Whether to do it turns on the operational layer, which is most of the product.
Procedure
Starting a Programme From Nothing
The first ninety days, in order. Almost none of it requires a purchase, and the sequence determines whether the programme is trusted or resisted.
Reference
Terms defined once, including the ones vendors and practitioners use to mean different things.
Reference
Terms used across these notes, defined once, including the several that are used inconsistently across the industry.
Software guides
Independent comparisons for teams choosing monitoring, time tracking, workforce analytics and contact-centre QA platforms.
Top 5
Employee Monitoring Tools for Contact Centre Teams
Five platforms compared through transparency, reporting and pilot controls.
Top 8
Time Tracking Tools for QA and Operations Teams
Eight options for calibration, coaching, reporting and project work.
Top 12
Workforce Analytics and QA Platforms
Twelve tools separated by data layer and operating model.
50 notes and 3 software guides