QA sampling in contact centres is sound, and that is why the imbalance goes unnoticed

In a global survey of more than 900 customer experience leaders, 63% pick review calls at random, one way to meet the standard's requirement that the method be unbiased. The sample is sound. What sits beside it is not: handle time, hold and adherence are computed by the phone systems on everything they touch, while quality is read by a person on a handful of calls a month. Spellit Actions scores the conversation side of every call.

Andrey TikhomirovCRO, Spellit8 min read

Reviewing every call instead of a sample changes less than it promises, and that argument is made elsewhere and still holds. This is the other half of it: uneven coverage is already deciding which of two conflicting goals wins.

The sample is not the problem

Review queues in sales teams fill with whatever was easiest to find, usually a complaint or an escalation. Contact centers mostly do not work that way. A global benchmarking survey published by COPC in May 2022, fielded from September to December 2021 among more than 900 customer experience leaders across the Americas, Europe, Africa, Asia and Oceania, found 63% selecting conversations at random, 25% letting the reviewer choose, and 12% doing something else. Ninety-two percent run a quality program at all.

Two things to hold about that source. COPC runs and paid for the survey, and it also sells consulting and certification against its own standard, so it has an interest in the subject being taken seriously. And the report gives no sample size per question, only the overall figure of more than 900, with no industry breakdown.

Even allowing for both, the diagnosis that works for a sales team does not transfer here. The structural problem is a different one, and it has nothing to do with which calls get picked. What does your own month of conversations actually contain?

Everything except quality is measured on every call

Coverage is the asymmetry nobody chose. Average handle time, hold time, after-call work, transfers and schedule adherence are computed by the telephony and workforce systems, which is why they cover everything those systems touch. Quality is read by a person on a handful of conversations per agent per month. Nobody publishes that handful as a figure with a sample behind it, so take four a month as an illustration and substitute your own. Two measures of the same work, one complete and one sampled, pointed at goals that pull in opposite directions.

ContactBabel's 2024 guide to the US contact center market, built on structured interviews with 189 US contact center directors and managers fielded between September and December 2023, makes the consequence visible and draws the conclusion itself, calling the two lists somewhat disconnected. Ranked by what centers say they encourage, customer satisfaction comes first and short handle time sixth. Ranked by what they actually reward, attendance comes first, adherence second, short handle time third, and customer satisfaction fourth.

The report publishes an ordering rather than weights, so the size of those gaps is unknown, and the two columns are ranked by different measures: top-three mentions on one side, greatly or somewhat rewarded on the other. The field work is also now around three years old. It is a sponsored publication with editorial independence stated in the document, and it covers the United States only. Read it as the shape of a problem rather than a measurement of its size. This is the tension a review of contact center conversations has to sit inside, and caring more about quality does not resolve it. There are only two ways out of an imbalance: raise the weaker measure or lower the stronger one. Nobody is going to stop computing handle time on every call, which leaves one option standing.

The mechanism needs no villain. A supervisor who listens carefully, calibrates monthly and runs honest one-on-ones still produces a quality number drawn from four conversations, next to an efficiency number drawn from four thousand. Nobody decided the second should outrank the first. The arithmetic decided.

The standard asks for three scores, and puts the highest bar where the check is mechanical

Quality in this market is not one number, and the governing document says so. The COPC standard for customer operations, release 7.0, written by an organization that also sells certification against it, requires that three kinds of accuracy be assessed as distinct components, each with its own target. It also requires that the method for picking the sample be unbiased, without saying it must be random.

What is being checkedTypical target in the standardWhat the check involves in practice, our reading
Compliance accuracy, against the regulator's requirements for that business99.5%, varying with what the regulator asksReading a checklist against a transcript
Customer accuracy: the query was solved, the person was treated well, the explanation was clear95%, or 98% where only satisfiers are measuredReading the conversation and judging it
Business accuracy, against the rules the operation or its client sets90%Mostly checklist, some judgment

Look at the direction of those bars. The highest one sits on the check a form can do unaided. The lowest sits on the question of whether the customer got what they called about. That ordering is defensible, since a regulatory miss carries a fine and a clumsy explanation does not. It also means a conversation can pass the form at a hundred percent and still have resolved nothing, and the scoring sheet will not report that as a failure.

The ordering also has a consequence for coverage, which is why it belongs in this piece. A compliance check is a checklist read against a transcript, so it costs almost nothing per conversation and could run on all of them. The customer-accuracy check needs a person to read the conversation and decide, and that is the one rationed to a handful per agent per month. The split between what gets measured completely and what gets sampled runs along the same line inside the quality score as it does outside it. The regulatory requirement itself is the client's own, set by whoever supervises their industry. How many of your calls ended with the question unresolved

The score mostly stays with the agent

What happens to a quality score after it is produced turns out to be narrow. In the 2024 US study of 189 contact center directors, 65% called their quality program very effective for assessing an agent and their training needs, and 42% called it very effective for regulatory audit. For knowledge the rest of the organization could use, 12% said very effective and 40% called it ineffective, which is the bottom of a three-point scale rather than the extreme of a five-point one.

So the instrument works as an audit and a coaching aid for one person at a time, and almost nothing of what it learns reaches product, marketing or operations. The same survey of 189 directors found 78% naming lack of time to analyze the data as a problem, and 40% calling it a serious one.

A handful of conversations a month per agent cannot support a statement about what customers are calling about and what goes wrong, however carefully each one is read. That is the coverage problem again, wearing a different coat. The conversations needed for that statement already exist and are already recorded, which is the narrow thing reading every call for quality is good for.

Insurance runs into a version of the same trap from the other side, where an inbound service call gets closed quickly because handle time is what the department is measured on. The conversation and the clock are scored by different instruments there too.

Three quarters of the list we built was noise

We assumed the interesting calls could be found without knowing what the operation is for. On a contact center project, 34 of 116 conversations turned out to be relevant to the thing the client actually wanted to improve; the rest were inquiries that were never going to become anything. Without a relevance filter, the list of conversations worth attention was three quarters noise, and a list that is three quarters noise gets closed and not reopened. We added the filter because of that project. It works, and it costs the client a judgment call about what counts as relevant, which is one we cannot make for them.

The denominator of that project's headline measure was set to all calls, not to the relevant ones, and that was a decision rather than an oversight: the point of the work was an honest picture of what customers were saying, and narrowing the base would have flattered it. The same screen in a sales team has to divide by something else entirely. We had been treating that denominator as a constant, and it is a setting. The practical rule we took from it: before anyone compares a rate across two teams, somebody has to be able to say out loud what each one divides by. If nobody can, the comparison has not happened yet.

What to do next

Count the two coverages side by side for one month: how many calls produced an automatic efficiency number, and how many produced a quality score. Write both figures on the same line. Whatever the ratio turns out to be on your floor, put it on the slide where the quality target is reviewed, because that is the meeting where the two numbers currently appear as though they were comparable.

Then take one agent's full set of scored conversations for the quarter and check how many of the customer-side questions, the ones about whether the query was solved, were answered from the recording rather than inferred. That tells you how much of the quality score is currently a checklist and how much is a reading of the conversation.

One of the two numbers is already yours. Your telephony has been counting the first one all along. Send us a month of calls and we will produce the second at the same coverage, so the ratio can finally be written on one line. We will bring the second number.
Key points
  • 63% of centers already pick review calls at random, which is what the standard asks for.
  • Operational metrics are computed by the telephony and workforce systems, so they cover everything those systems touch. Quality needs a person.
  • The same survey separates what centers encourage from what they reward, and the two lists disagree.
  • The quality standard asks for three different scores with three different bars, highest where the check is mechanical.
  • Quality scores rarely leave the agent. Only 12% call them very effective for knowledge the rest of the organization uses.
See the demo now. Then run it on yours.

A live demo of our products and a real conversation about the growth problem you need to solve.

Book a demo

FAQ

How many calls per agent should be reviewed each month?

There is no defensible published benchmark. The figures circulating, including the familiar four or five per agent, trace to vendor blogs without a sample or a period behind them. The better question is the ratio: how many calls produce an efficiency number versus a quality score. Once that ratio is written down, the number per agent stops being the interesting part.

Is random sampling the right way to choose calls for review?

For an audit, yes, and the standard requires an unbiased method. In a global survey, 63% of centers already do it. Random sampling is the correct answer to the question of whether the operation meets its targets. It is the wrong tool for finding out what customers are calling about, because a random handful cannot support that kind of statement.

Why do quality scores and customer satisfaction disagree?

Usually because they measure different things and one is mostly a checklist. The compliance portion of a quality score carries the highest target and is checked mechanically; the part about whether the customer got what they wanted carries a lower target and needs the conversation to be read. A call can clear the form completely and resolve nothing.

Does full coverage of calls fix quality on its own?

No. Coverage on its own moves nothing, and [the evidence for that has its own piece](/blog/playbooks/sales-call-review-process/). What full coverage fixes here is narrower and specific to this market: it lets the conversation side be measured as completely as handle time, so the two goals can be argued about on equal terms instead of one winning by arithmetic.

What should a supervisor stop doing to make room for this?

Nothing about listening, calibrating or running one-on-ones, all of which are the parts that work. The thing to stop is treating a monthly score drawn from four conversations as a statement about the agent's month. It is an audit sample, and the questions about what customers are saying belong to the full set of recordings instead.

How does agent turnover affect any of this?

It shortens the horizon of everything built on individual scores. In a 2024 survey of 189 US contact center directors, attrition averaged 31% with a median of 24%, and 12% of operations lose more than a quarter of new starters within three months. A monthly cycle of scoring and feedback assumes the person is still there next quarter, which for a meaningful share of a floor is not a safe assumption.