Skip to content
QASignal Room

Notes  /  Operations

Evaluator Workload and Score Quality

Score quality degrades measurably with volume and session length. Most programmes set evaluator targets without knowing where that threshold sits.

Section
Operations
Type
Analysis

Evaluation is cognitively demanding work performed to a quota. The quota determines the quality of every score the programme produces.

What degrades with volume

Attention. Later evaluations in a session are scored differently from earlier ones, measurably.

Detail. Under time pressure, evaluators score the items that are quick to check and skim the ones requiring judgement.

Consistency. Agreement with other evaluators falls as volume rises.

Comment quality. The written feedback, which is what the agent actually reads, becomes shorter and more generic.

Anchoring. After several poor calls, an average call scores higher than it would in isolation.

Establishing your own threshold

Testable, and almost nobody tests it.

Insert a known calibration call at several points in an evaluator's session — early, middle, late.

Compare the scores. If late-session scores on the same call differ systematically from early ones, the session is too long.

Do the same across a week to find the volume effect.

Set the quota below the point where drift appears, not at the point where evaluators can just about complete it.

What a realistic quota looks like

A thorough evaluation of a ten-minute call takes twenty to thirty minutes — listening, scoring, writing feedback. Faster than that is skimming.

Which puts a full day at perhaps twelve to sixteen evaluations, before any other duty.

Most programmes set quotas well above this and receive scores produced accordingly.

The arithmetic is worth doing openly: quota multiplied by realistic time per evaluation, compared against available hours. Where it does not fit, the scores are being produced by a process nobody would defend if described.

How analytics changes the calculation

Not by making evaluators faster at the same task, which is the assumption in most business cases.

By changing which calls they evaluate. A targeted queue of calls that actually need judgement is more valuable per evaluation than a random sample, so fewer evaluations produce more insight.

By removing the mechanical checks. Presence of a disclosure does not need a human, which frees the human for the parts that do.

The correct response to analytics is usually fewer, deeper human evaluations, not the same number faster. Most deployments do the opposite and increase the quota.

The evaluator role

It is a specialist role requiring product knowledge, judgement and the ability to write feedback someone will accept.

It is frequently staffed by rotation from the floor, which brings product knowledge and no evaluation training.

It is isolating. Listening to calls alone all day, delivering criticism, with limited social contact.

Turnover is high in many operations, which resets calibration continuously.

Mitigations that work: part-time evaluation combined with other duties, rotation, calibration sessions that double as team contact, and involvement in the process findings so the work connects to visible improvement.

The measure worth watching

Agreement over time, from calibration.

If agreement is falling, something is wrong — volume, turnover, scorecard drift, or fatigue.

Track it alongside evaluator volume. The relationship between the two is usually visible and it is the argument for the quota you want.

A programme that cannot say how many evaluations per day it expects, or what its inter-rater agreement is, does not know the quality of its own primary output.

The embedded calibration call

A cheap instrument for detecting fatigue and drift within a session.

Choose one call with an agreed score.

Insert it into evaluators' queues unannounced, at varying positions — second, eighth, fifteenth.

Compare the scores by position. A systematic difference between early and late is session fatigue and it sets your maximum session length.

Compare across weeks. Drift over time is calibration decay.

Compare across evaluators, which is the standard calibration use.

Rotate the call periodically so it is not recognised.

This is the only way to measure fatigue effects, it costs one evaluation slot per session, and it produces the evidence for the quota you want.