Skip to content
QASignal Room

Notes  /  Buying

Evaluating a Speech Analytics Platform

Demonstrations run on clean audio with tuned categories. The questions and tests that reveal what the product does on your calls.

Section
Buying
Type
Checklist

Every platform demonstrates well, because the demonstration uses the vendor's audio and categories built for it.

Test it on your calls

Non-negotiable. A pilot on your own recordings, including your worst audio.

Include your accent range, your queues, your product vocabulary.

Include your mono recordings if you have any, and your noisy environments.

Measure word error rate yourself against a human transcript on thirty calls, stratified by group. Do not accept the vendor's figure.

Build three categories from your own operation and measure their precision on fifty calls each.

The questions

What is your word error rate on telephony audio in our accent mix? A vendor who has tested will discuss it; one who quotes a benchmark figure has not.

How is transcription accuracy affected by accent, and do you measure it by group? The answer indicates whether they have engaged with the problem at all.

What is rule-based and what is model-based in categorisation and scoring?

Can we see and edit the category definitions? A black box category set cannot be maintained.

Can we pin a model version, and are we notified of updates?

How are automated compliance determinations explained to an agent disputing one?

Does it support stereo capture, and what degrades without it?

What acoustic measures are available independent of transcription?

Can emotion inference be disabled as a component?

Can we export everything — audio, transcripts, determinations, categories — in a usable format?

What to weight lightly

Dashboard aesthetics, which is where demonstrations spend their time.

Number of out-of-the-box categories, which are generic and will be replaced.

Accuracy percentages without a stated test set and method.

Emotion and sentiment capability, for the reasons in their own notes.

Analyst rankings, which reflect market presence.

The pilot design

Sixty days minimum. Long enough to build categories, measure them, and see whether anyone uses the output.

Your people operating it, not the vendor's.

A defined question to answer: can this detect our compliance requirements at a precision we can act on, on our audio.

Success criteria written before the pilot, with numbers.

Measure your own hours. Integration and category building are the underestimated cost.

What the pilot usually reveals

Accuracy on your audio is lower than the brochure.

Categories need substantially more work than the demonstration suggested.

Stereo capture matters more than expected, and you may not have it.

The vendor's default category set is not usable and building your own is the real project.

The dashboard is not where the value is; the search and the exports are.

The decision

Buy on: accuracy on your audio, category tooling, explainability, export, and the data terms.

Not on: the interface, the category count, or the sentiment feature.

And budget for the category building, which is the ongoing work that determines whether the platform produces anything. A licence with nobody maintaining categories is an expensive transcription service.

The scoring sheet

Criteria weighted before the first demonstration, so the evaluation cannot be steered by the interface.

Transcription accuracy on your own audio, by speaker group. The highest weight, because it caps everything else.

Category tooling: can you build, see, edit and measure your own definitions.

Explainability of any automated determination reaching a person.

Export completeness, tested during the trial.

Data terms: training, sub-processors, retention, deletion, region.

Acoustic measures available independently of transcription.

Stereo support and degradation without it.

Model version control and update notification.

Integration effort, measured in your own hours during the pilot.

Weight these before you see a demonstration, and hold to the result. Evaluations drift toward whichever product presented best unless the criteria were fixed first.

External reference: NIST model-risk guidance.