Speaker Separation and Why Stereo Matters
Knowing who said what is the foundation of every agent-level measure. Mono recording makes it a guess, and a large number of deployments discover this late.
Almost every useful output depends on attributing speech to the right speaker. How that attribution is achieved determines its reliability.
The two approaches
Stereo capture. Agent on one channel, customer on the other, separated at the telephony layer. Attribution is exact.
Diarisation. Both speakers on one channel, separated afterwards by algorithm clustering voice characteristics.
The difference in reliability is large, and it is not a close call.
What depends on it
Every agent-level measure. Did the agent say the disclosure, or did the customer say something that sounded like it.
Talk ratio and talk-over, which are undefined without channel separation.
Sentiment attribution. Customer sentiment and agent sentiment are different measures and conflating them is meaningless.
Compliance determinations, which must attribute a phrase to the agent.
Coaching, which requires knowing which utterances were the agent's.
Where diarisation fails
Similar voices. Same gender, similar pitch, similar accent.
Crosstalk, which is exactly when the conversation is most interesting.
Short utterances. Backchannel — "mm-hmm", "right", "okay" — is frequently misattributed and it is a large share of turns.
Three-party calls, including transfers and supervisor joins.
Poor audio, which is the general case in telephony.
Speaker changes at low energy, such as the start of a turn.
Error rates on conversational telephony audio are not small, and every downstream attribution inherits them.
Establishing what you have
Surprisingly often unknown.
Check whether your recordings are stereo with agent and customer separated. Not whether the file is stereo — whether the channels are meaningfully separated.
Check the whole estate. Different queues, different telephony paths and different sites frequently differ, and a mixed estate produces analytics that work well for some queues and poorly for others.
Check what happens on transfers and conferences.
This is a five-minute investigation and it determines what the platform can do.
Fixing it
Stereo capture is a telephony configuration, not a software purchase.
It is the single highest-return change available to a speech analytics deployment, ahead of any model or platform improvement.
It may require work on the recording platform or the carrier configuration, and it is worth doing before rather than after buying analytics.
Where it genuinely cannot be done, know that agent-level attribution is approximate and do not build compliance determinations or performance measures on it.
Measuring diarisation accuracy
Where you are stuck with mono, measure it.
Take thirty calls, transcribe with diarisation, and check attribution by hand.
Count misattributed turns, and note whether errors cluster on certain speaker pairs.
Report the error rate alongside any agent-level metric derived from mono audio.
A compliance determination from mono audio with a 10 percent attribution error rate is a determination that is wrong about the speaker one time in ten, and that should be visible to anyone acting on it.
The purchasing implication
Ask the vendor what degrades without stereo, specifically, feature by feature.
Ask whether their pricing or accuracy claims assume stereo. They generally do.
If your estate is mono, pilot on mono. A pilot run on stereo sample audio proves nothing about what you will get.
The five-minute estate check
Before any analytics decision, establish what your recordings actually are.
Pull ten recordings from each queue and each telephony path.
Open them in any audio tool and look at the two channels.
Genuinely separated means agent audio on one channel and customer on the other, with near-silence on the opposite channel during each speaker's turn.
A stereo file with identical content on both channels is mono in every way that matters.
Check transfers and conference calls, which frequently collapse to mono even where the main path is stereo.
Check each site. Estates acquired or migrated at different times differ.
Record the finding per queue. It determines which analytics features are reliable where, and it is the input to every subsequent decision.
More in this section
- Transcription Accuracy and What Degrades It
- The Accent and Dialect Accuracy Gap
- Categorisation and Topic Detection
- Sentiment Analysis: What It Measures
- Emotion Detection and the Science Problem
- Acoustic Measures: Silence, Talk-Over and Pace
- Redaction, PCI and Sensitive Data in Recordings
- Real-Time Analytics and Agent Assist