Mechanics
Under the transcript
Transcription accuracy sets the ceiling for everything downstream, and it varies systematically by speaker. This section covers what each component actually does and where it fails.
Reference
Transcription Accuracy and What Degrades It
Vendor accuracy figures come from clean benchmark audio. Contact centre calls are not that, and the difference is large enough to change what the system can support.
Analysis
The Accent and Dialect Accuracy Gap
Speech recognition performs measurably worse for some speakers, along lines that map onto race and region. That has consequences.
Procedure
Categorisation and Topic Detection
The most useful thing speech analytics does, and the thing most often set up badly. Category definitions are the whole product and they need maintenance.
Analysis
Sentiment Analysis: What It Measures
It classifies language as positive or negative. That is not the same as knowing how the customer felt, and the difference matters.
Analysis
Emotion Detection and the Science Problem
Products claim to infer emotional state from voice. The scientific foundation is contested and regulators have begun restricting it.
Reference
Acoustic Measures: Silence, Talk-Over and Pace
The most reliable outputs in the whole toolkit, because they are measured rather than inferred, and the least discussed.
Procedure
Redaction, PCI and Sensitive Data in Recordings
Recordings capture card numbers, health data and identifiers. Redaction is a control with known failure modes, and so is pause-and-resume.
Analysis
Real-Time Analytics and Agent Assist
Prompting during a live call is a different product with different failure modes. When it helps, when it distracts, and the surveillance question it raises.
Analysis
Language Models in Conversation Analysis
Vendors have replaced rule engines with language models. This genuinely improves some tasks, introduces new failure modes, and makes the system harder to audit.
Reference
Keyword Spotting and Full Transcription
Two architectures with different economics and a decisive practical difference: whether you can ask a new question of last year's calls.
Reference
Speaker Separation and Why Stereo Matters
Knowing who said what is the foundation of every agent-level measure. Mono recording makes it a guess, and a large number of deployments discover this late.
11 notes in this section