Skip to content
QASignal Room

Notes  /  Programme

Defining Quality Before Measuring It

Most scorecards encode assumptions about what good looks like that nobody has tested. Deriving them from outcomes takes a quarter and changes what gets measured.

Section
Programme
Type
Procedure

Scorecards are usually assembled from industry templates, previous employers' forms and incident responses. They encode a theory of quality nobody has examined.

The question to answer first

What outcome are we trying to produce?

Resolution without repeat contact. A customer who stays. A sale. A regulatory obligation met. Reduced handle time without harming the first three.

Different answers produce different scorecards, and an operation that has not chosen will build a form that pursues all of them and optimises none.

Deriving items from outcomes

The exercise that most programmes skip and that changes the form substantially.

One: pick the outcome. Repeat contact within seven days is a good starting point because it is objective, available, and matters commercially.

Two: take a sample of calls with the outcome and without it. Several hundred of each.

Three: score them all on a broad set of candidate behaviours, including everything currently on your scorecard and anything else plausible.

Four: compare. Which behaviours differ between the two groups?

Five: keep those. Drop the rest.

The result is usually uncomfortable. Items that everyone believed mattered show no relationship, and something nobody scores turns out to separate the groups clearly.

What this typically finds

Patterns that recur across operations, though yours will differ:

Setting expectations explicitly — telling the customer what happens next, when, and what they need to do — relates to repeat contact strongly.

Confirming understanding before closing.

Diagnostic questioning early in the call.

Tone items frequently show weak relationships to hard outcomes, which does not mean they do not matter, and does mean they should not carry the weight they usually do.

Script adherence frequently shows no relationship at all, and sometimes a negative one.

The confounders

Call type. Complex calls both take longer and repeat more. Compare within a category.

Agent tenure, which correlates with everything.

Customer characteristics. Some customers repeat regardless.

Reverse causation. A difficult call causes both the behaviour and the outcome.

None of these invalidates the exercise; they mean the finding is an association to investigate rather than a proof. It is still enormously better than a form assembled from assumptions.

Doing it without a data team

A simplified version is available to anyone.

Take fifty calls that resulted in a repeat contact and fifty that did not, matched roughly on call type.

Listen to them in mixed order, blind to which group.

Note what differs.

This is qualitative and it is a real finding. Supervisors who do this describe it as the most informative week they have spent, and it costs nothing but time.

Rebuilding the scorecard

Keep the compliance section separate, as described elsewhere.

Behavioural items limited to those with an evidenced relationship to an outcome you named.

Weight by strength of relationship, roughly, rather than by argument.

Re-run the analysis annually. What predicts outcomes changes as the operation, the products and the customers change.

The organisational obstacle

The exercise frequently shows that some long-standing items do nothing, and those items have owners.

Present it as evidence, not as criticism. The form was built on the best available thinking at the time; now there is data.

Expect resistance to removing items, particularly ones added after an incident. The compromise that works is moving them to a compliance checklist rather than the scored section, which preserves the control without diluting the measurement.

Running the exercise with a small team

The full outcome analysis needs data support. A version that does not.

Pull one hundred calls with a repeat contact within seven days, and one hundred without, matched roughly on category.

Shuffle them. The listener must not know which group a call is in.

Two people listen independently to fifty each, noting what they observe rather than scoring.

Then reveal the groups and look at what the notes say about each.

The differences are your candidate scorecard items.

A week of two people's time, no analytics required, and the output is a scorecard grounded in something rather than assembled from a template.

Repeat with a different outcome — escalation, complaint, sale — and compare. Items that appear against several outcomes are the ones to weight.