Skip to content
QASignal Room

Notes  /  Programme

Calibration: Keeping Evaluators Consistent

Two evaluators scoring the same call disagree more than anyone expects. Measuring the disagreement is the only way to know whether your scores mean anything.

Section
Programme
Type
Procedure

If two evaluators score the same call differently, the score describes the evaluator as much as the agent. Most programmes have never measured how much.

Measuring agreement

Take one call. Have every evaluator score it independently, blind.

Compare the totals and, more usefully, each item.

Expect wide disagreement on the first attempt. Totals varying by fifteen or twenty points across evaluators is common, and it is a shock to programmes that have never checked.

The item-level view is where the fix is. Aggregate agreement hides which items are the problem, and it is usually two or three.

Repeat monthly. Agreement drifts.

What causes disagreement

Vague items. "Demonstrated empathy" produces the widest spread of any item on any scorecard, reliably.

Missing criteria for score points. A five-point scale with no anchors is five evaluators' opinions.

Undefined applicability. An item that does not apply to some calls, with no rule about what to do, is handled differently by each evaluator.

Different tolerance. Some evaluators mark harder. This is a person-level effect and it is measurable.

Knowledge of the agent. Evaluators score agents they know differently, in both directions.

Fatigue. Scores drift through a session.

Fixing it

Rewrite the disagreeing items, with observable criteria and anchored score points. This addresses most of it.

Add examples. A short library of scored calls illustrating each score point on each item is the single most effective calibration tool and it takes a day to build.

Blind the evaluator to the agent where the system allows. This removes a bias that is otherwise unfixable and it is technically straightforward.

Rotate evaluators across teams, so no evaluator consistently scores the same agents.

Cap session length. Fatigue effects are real.

Track evaluator-level averages. An evaluator consistently a few points below the group is applying a different standard, which is a conversation rather than a fault.

Running a calibration session

Monthly, an hour, everyone who scores.

Two or three calls, scored independently before the session.

Reveal the scores together. Discuss the items with the widest spread.

The output is a decision: either the item is rewritten, or a convention is agreed and documented.

Record the conventions. A calibration session that produces agreement in the room and no written record has to be repeated.

Measure agreement each time and track it. Improving agreement over months is the evidence the programme is maturing.

The uncomfortable finding

Many programmes that measure agreement for the first time discover their scores carry less information than assumed.

The response should not be to stop measuring agreement. It should be to fix the scorecard, because the disagreement was always there and was simply invisible.

Until agreement is reasonable, agent-level comparisons are not defensible. Ranking agents on scores from evaluators who disagree with each other ranks the evaluator assignment as much as the agents.

Calibrating automated scoring

The same discipline applies where a system scores.

Compare the system's determination against a human panel on a set of calls.

Measure agreement per item, as with human evaluators.

Check agreement by speaker group, which is where automated scoring can be systematically unequal.

An automated scorer that disagrees with the human panel is not automatically wrong — it may be more consistent — and the disagreement has to be understood before either is trusted.

The scored example library

The single most effective calibration tool, and it takes a day to build.

For each scorecard item, select two or three calls that clearly demonstrate each score point.

Agree the score in a calibration session, with the reasoning written down.

Store the clip, the score and the reasoning where evaluators can reach it during evaluation.

Reference it in disputes. "This is what a three looks like on this item" resolves most disagreements immediately.

Use it for onboarding. A new evaluator working through the library reaches usable agreement far faster than one learning by correction.

Refresh it when the scorecard changes, and retire examples that no longer reflect current practice.

Programmes with a library have measurably better agreement than those relying on written criteria alone, because the criteria are always more ambiguous than their authors believe.