Skip to content
QASignal Room

Notes  /  Foundations

Manual and Automated Evaluation: What Each Catches

They fail in opposite directions. Automation misses meaning and covers everything; humans understand meaning and see almost nothing.

Section
Foundations
Type
Analysis

The choice is not between them. It is about which questions each is asked, and getting that division wrong produces either an unreviewable queue or a compliance system with opinions about tone.

What automated evaluation does well

Coverage. Every interaction, not a sample. This eliminates the sampling problem entirely for anything it can assess.

Presence and absence. Was a phrase said. Was a step taken. Binary, verifiable, auditable.

Counting and trending. How many calls about a topic, this week against last.

Acoustic measures. Silence, talk-over, monologue length. Measured rather than inferred.

Consistency. The same rule applied identically to every call, with no fatigue and no drift.

Finding needles. The twelve calls out of forty thousand containing something specific.

What it does poorly

Whether the answer was correct. A confident, well-structured, entirely wrong answer scores well on every automated measure.

Whether the resolution was appropriate to the customer's actual situation.

Judgement calls, which is most of what quality means.

Anything depending on words the recogniser missed.

Context. The same phrase is appropriate in one call and not in another.

Novel problems. A rule catches what it was written for.

What humans do well

Meaning. Was this call handled well, all things considered.

Correctness of the advice given.

Novel observations. A human notices something nobody wrote a rule for, which is how new categories get discovered.

Context and proportion.

Coaching input that an agent will accept, because it comes with an explanation.

What humans do poorly

Coverage. A few calls per agent per month.

Consistency. Evaluators disagree with each other and with themselves, which is why calibration exists.

Fatigue. Score quality degrades through a session.

Rare event detection. Something happening on one percent of calls is invisible in a four-call sample.

Being unbiased. Knowledge of the agent affects the score, which is measurable and uncomfortable.

The division that works

Automated for coverage and detection. Compliance presence checks across all calls. Category and topic counts. Acoustic outliers. Flagging calls for review.

Human for judgement, on calls selected by the automated layer rather than at random.

This inverts the standard practice. Most programmes review random calls and use analytics for reporting. The productive arrangement uses analytics to choose which calls a human should look at, which makes each human review far more valuable.

Selecting calls for human review

Once analytics is providing coverage, random sampling stops making sense.

Calls with poor outcomes — repeat contact, escalation, low survey score.

Calls flagged by a compliance rule.

Acoustic outliers — very long silences, heavy talk-over, unusually long calls.

A category under investigation.

Plus a random sample, small, to catch what the rules do not, and to detect whether the selection itself is missing things.

That last one matters. A review queue driven entirely by rules only ever sees what the rules find, and the blind spot is invisible from inside.

What not to automate

Anything feeding performance management or pay, without a human decision in the loop and a route to dispute.

Compliance failures with consequences, without human confirmation. An automated compliance failure based on a transcript error is a serious matter for the agent.

Anything where the accuracy gap by speaker group has not been measured.

The rule is simple: automation can decide where to look; a person should decide what it means whenever the answer affects someone.

Designing the review queue

Once analytics provides coverage, the human review queue becomes a design decision rather than a random draw.

Compliance flags, where an automated check found a possible failure. Highest priority because the consequence is largest.

Outcome-selected calls: repeat contact, escalation, complaint, low survey score.

Acoustic outliers: extreme silence, heavy talk-over, unusual duration.

Category-driven: calls in a topic currently under investigation.

New agents, proportionally more.

A random sample, small but present, which is the only way to detect what the rules are missing.

Cap the queue at reviewable volume. A queue larger than capacity is triaged by whoever opens it first, which reintroduces selection bias in a less visible form.

Record why each call entered the queue, so the composition can be reviewed and adjusted.