Detecting Scorecard and Model Drift
Scores move for reasons unrelated to performance: evaluator turnover, model updates, call mix. Instrumenting distinguishes them.
A quality score changes and everyone assumes performance changed. Several other things move it, and they are distinguishable.
What causes drift
Evaluator turnover. New evaluators score differently until calibrated. A team's score can shift on a staffing change.
Calibration decay. Without regular sessions, evaluators diverge over months.
Scorecard changes, obviously, and the discontinuity is frequently not annotated.
Call mix changes. A new campaign, a product launch or a seasonal shift changes what agents are handling.
Model updates, where automated scoring is used. The vendor updates and the determinations change with no notice.
Transcription changes. A new recogniser version alters what is detected.
Audio infrastructure changes. New headsets, a codec change, a telephony migration all move accuracy and therefore every transcript-derived measure.
Instrumenting for it
A fixed calibration set. Ten calls, scored quarterly by all evaluators and by the automated system. The same calls, indefinitely.
Score movement on the fixed set is drift by definition, because the calls did not change. This is the cleanest possible detection and it costs an hour a quarter.
Track inter-rater agreement over time.
Track evaluator-level averages, to detect a new evaluator scoring differently.
Track call mix — category distribution, average duration — alongside scores.
Log every change: scorecard revisions, model updates, evaluator changes, telephony changes, with dates.
The change log
The single most useful artefact for interpreting a score movement.
When a score moves, check the log before the operation.
Most unexplained movements resolve there, and the ones that do not are the real findings.
Vendors update models without announcing it. Ask for notification in the contract, and detect it independently with the fixed calibration set, because notification is unreliable.
Automated scoring drift specifically
A model update can change determinations on identical audio.
Historical comparisons break across the update.
Compliance determinations made under a previous version cannot be reproduced, which is a problem if one is ever challenged.
Mitigations: version every determination, keep the audio and the transcript, run the fixed calibration set after every update, and require notice of updates contractually.
Presenting a movement
Never report a score change without checking the log.
When reporting, state what was ruled out: no scorecard change, no evaluator turnover, calibration stable, call mix stable, no model update.
A movement with those five ruled out is a finding. Without them it is a number that moved.
This is the same discipline as ruling out confounders in any analysis, and it is skipped for the same reason: the interesting explanation is more appealing than the mundane one.
The annual review
Compare this year's calibration set scores to last year's. Cumulative drift is visible here and nowhere else.
Compare the scorecard to the version a year ago, item by item.
Compare the evaluator pool. How much turnover.
Then ask whether any trend reported over the year survives these adjustments. Frequently a two-point improvement dissolves entirely, and knowing that is better than continuing to report it.
The fixed calibration set
Ten calls, unchanging, scored quarterly by everyone and by the automated system. The cleanest drift detector available.
Select for range: easy, difficult, ambiguous, compliance failure, technically complex.
Agree the reference scores once, in a full calibration session, and record the reasoning.
Re-score quarterly without reference to the previous results.
Any movement is drift, because the calls did not change. There is no other explanation to rule out, which is what makes this instrument so useful.
Plot it over years. Cumulative drift is invisible quarter to quarter and obvious over a two-year series.
Run it immediately after any model update to detect vendor changes independently of whether you were notified.
Refresh the set every few years, overlapping so the series is not broken.