Automated Scoring: Where It Works
Scoring every interaction is the headline promise. It works for a narrow band of items and produces disputes everywhere else.
The proposition is compelling: every call scored, no sampling problem, complete coverage. The reality is that only some scorecard items can be automated defensibly.
Items that automate well
Presence of a required phrase. Did the agent give the disclosure. Checkable, auditable, and the strongest case.
Absence of a prohibited phrase. Guarantees, unapproved claims.
Sequence. Was identity verified before account details were discussed.
Acoustic thresholds. Excessive silence, heavy talk-over, hold exceeding a limit.
Structural elements. Did the call include a summary at the end.
These share a property: the criterion is observable in the record and does not require interpretation.
Items that do not
Whether the information given was correct. The system does not know the right answer.
Whether the resolution suited the customer's situation.
Tone, empathy, rapport. Attempts to automate these produce scores agents dispute successfully.
Anything requiring context about the customer, the history or the product.
Judgement of appropriateness, which is most of what quality means.
The accuracy requirement
An automated item that reaches an agent's record needs to be very accurate, not merely better than chance.
Measure precision and recall per item, against a human panel, on your own calls.
For anything with a consequence, precision matters most. A false failure is an agent penalised for something they did.
Set a threshold and hold to it. Below it, the item runs in monitoring mode only and is not scored.
Check by speaker group. An item with good overall accuracy and poor accuracy for some accents produces unequal outcomes, which is the recurring theme of automated evaluation.
The dispute mechanism
Non-negotiable for automated scoring.
The agent can see which call, which item, and why.
They can hear the audio and read the transcript, side by side.
They can dispute, and a human reviews.
Disputes are logged and analysed. A high dispute rate on one item means the item is wrong. A pattern of upheld disputes from particular agents may indicate a recognition gap.
This mechanism generates the data that proves whether automated scoring is working, which is why programmes without it have no idea.
What automated coverage actually changes
The value is less about scoring and more about what full coverage enables.
Every compliance failure found, not a sample. This is the genuine transformation.
Trend visibility. A behaviour declining across the operation is visible immediately.
Targeted human review. Analytics selects the calls a human should evaluate.
Process findings. Aggregate patterns pointing at systems, knowledge and product.
The mistake is using it to produce an agent percentage to sit alongside the manual one. That combines the weaknesses of both.
A workable design
Compliance items automated across all interactions, reported as exceptions, confirmed by a human before any consequence.
Acoustic measures across all interactions, reported at queue and process level.
Behavioural items evaluated by humans on a sample selected by analytics.
No blended percentage. Two separate outputs, reported to different people, used for different things.
The agent sees both and can dispute either.
The precision threshold, set explicitly
Automated items need a stated accuracy bar and a rule about what happens below it.
Measure precision per item against a human panel, on at least a hundred flagged calls.
Set the bar by consequence. An item feeding a disciplinary process needs a much higher bar than one feeding a dashboard.
Below the bar, the item runs in monitoring mode: counted in aggregate, not attached to any individual.
Above the bar, it can be scored, with the dispute route active.
Re-measure quarterly and after any model change, because precision drifts.
Publish the precision figure to the agents subject to the item. An agent who knows an automated check is right 94 percent of the time understands why the dispute route exists and uses it appropriately.
An automated item with no measured precision should not reach anyone's record, and this is the most commonly violated rule in the field.
External reference: NIST AI Risk Management Framework.