Defining Quality Before Measuring It
Most scorecards encode assumptions about what good looks like that nobody has tested. Deriving them from outcomes takes a quarter and changes what gets measured.
Scorecards are usually assembled from industry templates, previous employers' forms and incident responses. They encode a theory of quality nobody has examined.
The question to answer first
What outcome are we trying to produce?
Resolution without repeat contact. A customer who stays. A sale. A regulatory obligation met. Reduced handle time without harming the first three.
Different answers produce different scorecards, and an operation that has not chosen will build a form that pursues all of them and optimises none.
Deriving items from outcomes
The exercise that most programmes skip and that changes the form substantially.
One: pick the outcome. Repeat contact within seven days is a good starting point because it is objective, available, and matters commercially.
Two: take a sample of calls with the outcome and without it. Several hundred of each.
Three: score them all on a broad set of candidate behaviours, including everything currently on your scorecard and anything else plausible.
Four: compare. Which behaviours differ between the two groups?
Five: keep those. Drop the rest.
The result is usually uncomfortable. Items that everyone believed mattered show no relationship, and something nobody scores turns out to separate the groups clearly.
What this typically finds
Patterns that recur across operations, though yours will differ:
Setting expectations explicitly — telling the customer what happens next, when, and what they need to do — relates to repeat contact strongly.
Confirming understanding before closing.
Diagnostic questioning early in the call.
Tone items frequently show weak relationships to hard outcomes, which does not mean they do not matter, and does mean they should not carry the weight they usually do.
Script adherence frequently shows no relationship at all, and sometimes a negative one.
The confounders
Call type. Complex calls both take longer and repeat more. Compare within a category.
Agent tenure, which correlates with everything.
Customer characteristics. Some customers repeat regardless.
Reverse causation. A difficult call causes both the behaviour and the outcome.
None of these invalidates the exercise; they mean the finding is an association to investigate rather than a proof. It is still enormously better than a form assembled from assumptions.
Doing it without a data team
A simplified version is available to anyone.
Take fifty calls that resulted in a repeat contact and fifty that did not, matched roughly on call type.
Listen to them in mixed order, blind to which group.
Note what differs.
This is qualitative and it is a real finding. Supervisors who do this describe it as the most informative week they have spent, and it costs nothing but time.
Rebuilding the scorecard
Keep the compliance section separate, as described elsewhere.
Behavioural items limited to those with an evidenced relationship to an outcome you named.
Weight by strength of relationship, roughly, rather than by argument.
Re-run the analysis annually. What predicts outcomes changes as the operation, the products and the customers change.
The organisational obstacle
The exercise frequently shows that some long-standing items do nothing, and those items have owners.
Present it as evidence, not as criticism. The form was built on the best available thinking at the time; now there is data.
Expect resistance to removing items, particularly ones added after an incident. The compromise that works is moving them to a compliance checklist rather than the scored section, which preserves the control without diluting the measurement.
Running the exercise with a small team
The full outcome analysis needs data support. A version that does not.
Pull one hundred calls with a repeat contact within seven days, and one hundred without, matched roughly on category.
Shuffle them. The listener must not know which group a call is in.
Two people listen independently to fifty each, noting what they observe rather than scoring.
Then reveal the groups and look at what the notes say about each.
The differences are your candidate scorecard items.
A week of two people's time, no analytics required, and the output is a scorecard grounded in something rather than assembled from a template.
Repeat with a different outcome — escalation, complaint, sale — and compare. Items that appear against several outcomes are the ones to weight.