Skip to content
QASignal Room

Notes  /  Programme

Multi-Language and Multi-Site Programmes

Analytics quality varies enormously by language, which means a global programme measures some sites better than others and reports the difference as performance.

Section
Programme
Type
Analysis

An operation running in six languages across four countries has six different levels of analytics capability and one dashboard.

The capability gap

Transcription accuracy differs substantially by language. Widely spoken languages with large training corpora perform well; others do not, and the gap is larger than the gap between accents within one language.

Category and sentiment models are usually built for one language first and ported. Ported models are weaker.

Some vendors support a language for transcription and not for downstream analysis, which is not always clear in the sales material.

The result: a site operating in a well-supported language appears to have better compliance and cleaner categories than one that does not, and the difference is the tooling.

What this does to reporting

Cross-site comparison is confounded by analytics quality.

A site with poorer transcription shows fewer detected compliance failures, which reads as better compliance.

Category volumes are not comparable across languages, because the categories perform differently.

Any global ranking measures the vendor's language coverage as much as the operations.

Handling it

Measure accuracy per language, on your own calls, as a first step. This is the same exercise described for accents and it produces the number every downstream comparison needs.

Report per language, not across. Each site against its own baseline and its own trend.

State the accuracy alongside the metric. A compliance rate from a language with 75 percent word accuracy carries a caveat.

Do not build a global league table. If leadership wants one, explain what it would be measuring.

Where a language is poorly supported, use manual QA more heavily and analytics for coverage of the few things it does reliably — acoustic measures work across languages, since they do not depend on transcription.

Translation

Several platforms offer translated transcripts so that a central team can review calls in languages they do not speak.

Useful for orientation and inadequate for evaluation.

Errors compound: recognition error, then translation error.

Nuance, politeness register and idiom are exactly what quality evaluation depends on and exactly what translation loses.

Do not score a call through a translation. Evaluation should be by someone who speaks the language, and a translated transcript is a tool for the analyst deciding which calls to send to them.

Cultural variation in what quality means

Beyond the tooling, the scorecard itself may not transfer.

Directness norms differ. Behaviour scored as confident in one market reads as rude in another.

Formality and address conventions vary and are frequently in the scorecard as a fixed requirement.

Silence tolerance differs measurably between cultures, which affects acoustic thresholds.

Rapport-building expectations vary widely.

A scorecard written for one market and applied globally measures conformity to that market's norms.

The workable arrangement: a common compliance section, a common structure, and locally adapted behavioural items with local calibration.

Multi-site within one language

Even without a language difference, sites differ.

Audio infrastructure varies, so transcription accuracy varies.

Call mix varies, so scores vary for reasons unrelated to quality.

Local calibration drifts apart without cross-site calibration sessions.

Run calibration across sites, not only within them, on the same calls, and measure the gap. It is usually larger than the gap between evaluators at one site, and it invalidates cross-site comparison until it is closed.

The per-language capability sheet

One page per language, maintained, that prevents most cross-site misreporting.

Measured word error rate on your own calls in that language.

Which analytics features are supported — transcription, categorisation, sentiment, redaction — and which are ported rather than native.

Category precision, measured locally rather than assumed from the primary language.

Whether acoustic measures are available, which they generally are since they do not depend on language.

Which evaluators are calibrated in that language.

The date of the last measurement.

Attach it to any report covering that site. A compliance rate accompanied by "measured word error rate 22 percent in this language" is read correctly; the same rate presented bare is read as a performance comparison.