Skip to content
QASignal Room

Notes  /  Foundations

What This Data Cannot Tell You

A list of questions routinely asked of quality and speech analytics data that it cannot answer, with what to use for each.

Section
Foundations
Type
Reference

A function that answers every question with the data it has loses the ability to be believed about the ones it can answer well.

Whether the advice given was correct

Why not: the system has no model of the right answer for the customer's situation.

Use instead: human evaluation by someone with product knowledge, on calls selected by analytics.

How the customer felt

Why not: sentiment classifies language and emotion inference has a contested scientific basis.

Use instead: what the customer said, quoted; survey responses; and acoustic measures described as what they are.

Whether an agent is good

Why not: small samples, evaluator disagreement, call mix differences, and scorecards that frequently do not relate to outcomes.

Use instead: outcome measures over long periods, aggregated, with call mix accounted for — and an acceptance that individual attribution is weak.

Why a customer contacted you

Why not: categorisation tells you what was discussed, which is not always the underlying reason. A billing call may be about a service failure.

Use instead: categorisation plus human review of a sample, plus the repeat contact reason.

Whether a process change worked

Why not: rarely, this is answerable — but only with a before period, an after period and a control, and most operations change several things at once.

Use instead: the same discipline as any measurement: baseline, single change, control group, adequate window.

Agent intent

Why not: a missed disclosure looks the same whether it was forgotten, skipped deliberately or said and misrecognised.

Use instead: the audio, a conversation with the agent, and the pattern across their other calls.

Anything about a specific individual from one interaction

Why not: one call is one call. The variation between an agent's calls exceeds the variation between agents.

Use instead: aggregation, over time, with the sample size stated.

Whether a customer will churn

Why not: this is sold and the evidence is weak. Sentiment and language features predict churn poorly compared to behavioural and account data.

Use instead: account data, tenure, product usage, prior complaint history, with call data as a marginal input.

The state of the operation from a single score

Why not: the score is compressed, weighted arbitrarily and gameable.

Use instead: compliance rate, outcome measures, process findings, and the programme health measures.

How to decline

Say what the data can support instead. "I cannot tell you whether the agent handled it well from analytics; I can tell you which calls a human should review and why."

Say what would answer it. Human evaluation, first-party outcome data, a designed experiment.

Say what it would cost.

Do not produce a number with a disclaimer for a question in this list. The disclaimer is dropped and the number is used, and when it is contradicted the whole function is discounted.

Offering the alternative

Declining a question is only useful when accompanied by a route to an answer.

For correctness of advice: human evaluation by a product expert on a targeted sample.

For customer feeling: survey free text, and quoted customer language rather than a sentiment score.

For agent capability: aggregated outcomes over a long period, with call mix accounted for.

For process effect: a designed comparison — baseline, single change, control group, adequate window.

For churn prediction: account and behavioural data, with call data as a marginal input.

For market or competitive questions: not this data at all.

And where nothing will answer it, say so plainly and say why. A function that names the limits of its evidence is believed when it presents evidence, which is the whole reason the limits are worth stating.