Calibration sessions are measured by whether they happen, not by what they fix
Nine in ten contact centers run calibration, and the industry asks no further question: published benchmarks count whether sessions happen and whether managers like them. Nobody publishes how far two assessors scoring the same conversation land apart. Where that has been measured, in an adjacent field that scores recorded calls against a form, trained assessors gave the identical score under half the time. Spellit Actions reads the same calls as a third scorer.
Everything measured below was measured on medical triage calls, not on sales calls. The mechanism is the same, a recording scored by trained people against a form of around two dozen items. I am going to argue that it transfers anyway, and say where I think it does not.
The industry measures whether calibration happens, not what it finds
Ask the benchmarks about calibration and you get attendance figures. A global survey of more than 900 customer experience leaders, fielded between September and December 2021, reports that 89% have a calibration process, 90% run it at least quarterly, and 86% consider their own calibration effective. Those are real numbers about a real practice.
Now read the same 39-page report for the words that would tell you whether calibration works: kappa, agreement, variance, correlation, standard deviation. Each appears zero times. The industry asks whether you calibrate and whether you like how you calibrate, and never asks how far apart your people land.
COPC ran and paid for that survey, and sells certification against the standard it measures against. That makes the next part more interesting rather than less: the standard behind the survey is stricter than the survey. It requires that everyone performing monitoring be calibrated at least quarterly using a quantitative approach that measures calibration at the level of the individual item, against a reference. Published practice has drifted a long way from published requirement, and what dropped out was the half that produces a number. The same gap shows up in what contact centers measure on every call and what they sample. One conversation scored twice by your own people shows where a form stops meaning one thing.
Where divergence has been measured, it is larger than anyone plans for
Two studies, two countries, one instrument and its translation, one answer.
A Dutch study published in 2017 took 114 recorded calls from four out-of-hours services and had each one scored by three of eight assessors, on a 24-item form with a five-point scale. The assessors were not casual: six certified triage staff, a physician and a site manager, averaging five and a half years of scoring experience. Complete agreement, meaning the identical score, came out at 44.9%. Allowing a one-point difference lifted it to 72.5%.
A Danish study published in 2019 translated the Dutch instrument and extended it, took 48 recorded contacts split between two designs, and landed in the same place: across the 24 scored by different assessors, average complete agreement was 0.40, ranging from 0.25 to 0.93 depending on the item.
Both were funded by bodies with an interest in quality improvement rather than in a particular result, and both are medical triage rather than commercial calls. These are the conditions everyone assumes are sufficient, trained people with years of practice and a written manual, and agreement still came out below half. Scoring conversations at all rests on the assumption that two readers of the same call would land in the same place.
The disagreement lives in the item, not in the people
Here is the finding that should change what a calibration session does. In both the Dutch and the Danish study, assessors were considerably more consistent with themselves than with each other. Rescoring the same calls four weeks later, the Dutch assessors matched their own earlier score 55.1% of the time against 44.9% with a colleague. The Danish study found the same split in reliability coefficients: good against oneself, poor against others.
Two people each applying a stable standard of their own are not undertrained. They are reading the same sentence to mean two things, and the per-item numbers show which sentence.
| What the item asked | Complete agreement |
|---|---|
| Did they ask to speak to the patient directly | 78.1% |
| Did they consult the doctor only when necessary | 62.3% |
| Did they check the caller understood the next step | 27.2% |
| Did they check the caller agreed with the next step | 23.7% |
| Did they pay attention to the caller's experience | 24.6% |
The items at the top describe an event that either occurred on the recording or did not. The items at the bottom ask for a judgment about quality of attention. Grouped, the medical items averaged 51.6% agreement and the communication items 39.3%.
The Danish study used a two-day course and a fuller manual, and reported that agreement between assessors stayed impaired anyway. So the instruction that follows is narrow: when two people disagree on an item repeatedly, rewrite the item. Do not schedule another session to align the people.
A coarser decision agrees better than a finer score
The Danish study also scored the same items as a binary, poor against sufficient. Average agreement rose from 0.40 to 0.75.
The same relationship sits inside the paper that gave the field its thresholds. Landis and Koch, in 1977, published the bands everyone still quotes, where 0.61 to 0.80 counts as substantial agreement. Their own worked example had two neurologists classifying the same patient records on a four-step scale; they agreed exactly on 43% of cases, and the coefficient came out at 0.21. Relaxing the requirement so that adjacent categories counted as agreement moved it to 0.60.
Two things follow. A five-point scale on a judgment item is buying precision the assessors cannot deliver, and a yes-or-no on an observable event is usually the better instrument. That choice is worth making before the form reaches the people who read the numbers it produces. If your form holds an item nobody has ever agreed on, an afternoon will find it. The bands themselves were published in 1977 as a way to talk about one worked example, which is a narrower job than the one they now do.
A model disagrees with your scorers about as often as they disagree with each other
The comparison nobody runs is model against human next to human against human. On a 2023 benchmark it came out at 85% and 81%. That work judged a pairwise preference between two answers rather than scoring one conversation against a form of items, so the number does not sit on the same scale as the 44.9% above and should not be subtracted from it.
What carries across is narrower and it is enough. A 2024 study of rubric-based evaluation, run at Microsoft and built on 741 ratings from 24 professional annotators, found model agreement with humans varying enormously by item: around 0.51 correlation on whether a citation was present, around 0.25 on efficiency. A model is not more consistent than your assessors. It is differently consistent, and the items where it cannot agree with anybody are the items to read out loud in the session. That is the same use a review of every call puts it to, one layer down.
We fixed two bugs and changed the client's history underneath them
A project changed the number of an item partway through its life, and our code treated that number as fixed for the whole history. Weeks that had been scored under the old definition were reported under the new one. Separately, a change in how a tag was spelled meant one recurring objection was counted as two, which pushed its share down.
We fixed both in August, and the fix itself changed the numbers on the client's screen retroactively. They had been comparing before against after, and the comparison had been measuring our bug.
What separates this from hardcoding a value that should have been a setting is that the thing fixed in place here was a moment in time. The error did not put a wrong number on today's screen. It rewrote what the previous quarter had said, after the fact, and we had no way to tell them which version of the past they had been reading.
What to do next
At your next calibration session, before any discussion, have everyone score the same two conversations independently and write down the spread per item. Not the average score, the spread. Items where two people land more than one step apart are the agenda; the rest is conversation.
Then take the three items with the widest spread and ask whether each one describes something that happened on the recording or something a person felt about it. Rewrite the second kind as the first, or collapse it to a yes-or-no. If neither is possible, the item is measuring the assessor rather than the call, and it should come off the form.
Before you send anything, have two of your own people score the same two calls separately. That alone will tell you something. Send the two recordings, the form and both filled sheets, and we will add a third sheet and hand back the spread across all three, item by item. Two calls, two of your scorers, and a third.
- The industry benchmark for calibration is attendance. We could not find a published measure of divergence in sales or contact centers.
- In two independent studies of scored call recordings, full agreement between assessors ran at 45% and 40%.
- Assessors were stable against themselves and unstable against each other. The disagreement lives in the wording of the item.
- Longer training did not raise agreement. Collapsing a five-point scale to a yes-or-no decision raised it from 0.40 to 0.75.
- The thresholds everyone quotes as an industry standard were described as arbitrary by the people who published them.
A live demo of our products and a real conversation about the growth problem you need to solve.
Book a demo