Skip to content
QASignal Room

Notes  /  Foundations

Anatomy of a Scorecard That Measures Something

Most scorecards accumulate items nobody removes, weight them arbitrarily, and produce a percentage that cannot be traced to any outcome. A shorter one works better.

Section
Foundations
Type
Procedure

Scorecards grow. Every incident adds an item, nothing is ever removed, and after three years the form has forty lines and takes twenty minutes to complete.

The structural problem

Items are added in response to events and removed never.

Weights are set by argument rather than by evidence about which behaviours affect outcomes.

Compliance and coaching items sit in the same total, so a missed disclosure can be offset by good tone. This is indefensible when a regulator asks.

The percentage is the output, and a percentage built from arbitrary weights over heterogeneous items is a number without a referent.

A better structure

Two sections, scored separately, never combined into one figure.

Section one: compliance. Pass or fail per item. No weighting, no partial credit. A call either contained the required disclosure or did not. The output is a list of failures, not a score.

Section two: quality. A small number of behavioural items, each with defined criteria, scored on a short scale. The output is a profile, not a single number.

Report them separately. "Two compliance failures this month; quality profile shows weakness on discovery questioning." That is actionable. "91.4%" is not.

Choosing the behavioural items

The hard part, and the part usually skipped.

Start from outcomes you care about: resolution, repeat contact, customer satisfaction, sales conversion, complaint rate.

Ask which agent behaviours plausibly affect them.

Then test. Score a sample, join to outcomes, and look at which items actually correlate. Items with no relationship to any outcome are measuring compliance with the form.

Most scorecards have never been tested this way and would lose half their items if they were.

Six to ten behavioural items is enough. More than that and evaluators cannot hold them in mind, calibration degrades, and the marginal items add noise rather than information.

Writing an item that can be scored consistently

Observable, not inferred. "Asked at least one open question to establish the reason for the call" can be verified. "Demonstrated empathy" cannot, and it produces the widest evaluator disagreement of any item on any scorecard.

Defined criteria for each score point, with examples. Not a five-point scale with no anchors.

Applicable to the interaction type. An item that does not apply to half of calls forces evaluators to invent a convention, and they each invent a different one.

Testable in calibration. If two evaluators consistently disagree on an item, the item is badly written, not the evaluators.

The empathy problem

Every scorecard has an item like this and it is the least reliable one.

It cannot be defined observably without reducing it to specific behaviours — acknowledging the customer's situation, using their name, not interrupting.

Reduce it to those behaviours, and score them. That is what the evaluator is actually looking for.

Or accept it is a judgement item and expect low inter-rater agreement, and do not weight it heavily.

What not to do is leave it undefined and weighted at 20 percent, which is the standard arrangement.

Reviewing the scorecard

Annually, with the evidence.

Which items have failed in the last year? An item nobody ever fails is not discriminating and should go.

Which items correlate with outcomes? Keep those.

Which items produce evaluator disagreement? Rewrite or remove.

Which items exist because of an incident three years ago that has not recurred?

Expect to remove a third of the form. The instinct is to add; the discipline is to cut, and a shorter scorecard is scored more consistently and completed more often.

Retiring an item

Scorecards grow because removal has no process. Giving it one makes the form maintainable.

Any item can be proposed for removal by an evaluator, a supervisor or an agent.

The test: has it failed in the last year, does it relate to any outcome, and do evaluators agree on it?

An item that never fails is not discriminating. Either the behaviour is universal, in which case stop scoring it, or the definition is too lenient.

An item with no outcome relationship stays only if it exists for compliance reasons, and then it belongs in the compliance section.

An item with poor evaluator agreement is rewritten or removed, not retained.

Record the removal and the reason, so it is not re-added after the next incident by someone who does not know it was tested.

Expect to remove items every year. A scorecard that only grows is a scorecard nobody is maintaining.