How to review every sales call, and why coverage on its own changes nothing

A sales call review process at full coverage scores every conversation automatically, then routes a small number of specific observations to the people who are not already doing the thing. Coverage is the easy half and moves little by itself: a 2023 meta-analysis of workplace monitoring finds no sign that monitoring improves performance. In Spellit Actions the score starts at 100 and every failed line takes points off it.

Oleg KulakovCEO, Spellit11 min read

Sales leaders listen to calls. They listen carefully, they take notes, and they hear things no review form will ever catch. That is the part most write-ups on this subject get wrong when they open with a line about managers who cannot be bothered.

The constraint was never effort. A week contains more conversations than one person can read, and the queue decides which ones get opened. A complaint pushes a call to the top. So does an escalation, a deal that visibly went sideways, a new starter whose first month somebody promised to watch. That ordering is the only one the week offers, and it is a perfectly sensible way to spend four hours. It is just not a sample. The arithmetic of that gap is in the coaching piece; this article is about what the conversations nobody opened have in common.

Whole categories of call never reach the shortlist

A queue built from complaints and escalations returns conversations that already announced themselves. The categories it skips are consistent, and each one holds a different kind of information. Opening more of the same calls does not close that gap.

Call that rarely gets reviewedWhy it never surfacesWhat it would have shown
A routine winNothing went wrong, so nobody flagged itWhat the team actually does when it works
A quiet loss with no complaint attachedNo escalation, no ticket, no noiseWhere deals leak without anybody objecting
A call that ended in four minutesToo short to look interestingWhether qualification is cutting off the right people
A strong rep having an ordinary dayThat name is not the one anybody is watchingWhether a problem has started spreading
A call in the team's second languageHarder to review, so it waitsWhether the script survives translation

Run the check on your own team before changing anything. Take the last twenty calls anyone reviewed and write, next to each, what put it in front of them. If the answers are mostly incidents and a handful of names, the process is producing an audit of a shortlist.

A statistically sized sample does fix the arithmetic, and it deserves a straight answer. The numbers below are ours, worked from the standard margin on a proportion; no survey was involved. Three hundred conversations pin a team-level rate to roughly five points either side. Scoring all thirty thousand of a contact center's monthly calls narrows that to about half a point. For the question "how often does this team do X", the other twenty-nine thousand seven hundred buy almost nothing.

The catch is which question the rate answers. A leader rarely needs the team average to a decimal place. They need to know whether this person, on these conversations, keeps doing one specific thing, and for that the useful sample is the whole of one rep's month: about thirty conversations in B2B. At that size the statistically adequate sample and the population are the same set, and the case for sampling quietly disappears. Bring us a week and we will tell you which rows of that table your own review has never opened.

Coverage on its own moves nothing

Reviewing every call does not by itself improve how anybody sells. A 2023 meta-analysis of electronic performance monitoring, pooling 94 independent samples and 23,461 people in total, reports no evidence that monitoring improves worker performance. Across the 41 samples that measured performance, covering 5,804 people, the pooled correlation is −.01 with an interval of [−.05, .04] sitting squarely across zero.

What monitoring does move is measurable and unhelpful. The same analysis puts the correlation with stress and strain at .16, interval [.13, .20], and with felt invasion of privacy at .28, interval [.13, .43] from only eight samples. Workers' attitudes at work drift negative overall, −.11 with an interval of [−.20, −.03]. Monitoring people in more ways at once was the one thing that significantly predicted more counterproductive behaviour, though on four samples and an interval that barely clears zero.

Two honest caveats, because this is the paper most likely to be quoted back at anyone selling this idea. Its performance samples are dominated by experiments, with field sales teams barely represented. And monitoring for developmental purposes, which is the case this article is about, is the one purpose the authors could not test against performance at all: they found two studies. What the paper does kill is the claim that watching everything makes people sell better, and it kills it thoroughly, because it separately analyses monitoring that is recorded and reviewed later and finds the same nothing there.

So before switching coverage on, write down who receives what, in what form, on which day. If that sentence does not exist yet, full coverage will produce a screen nobody opens, and the honest reading of the evidence above is that the screen will change no behaviour while raising everybody's stress.

A line with no quote behind it should come back blank

Every scored line needs four things attached: what was being checked, the verdict, the words the rep actually said, and where in the recording they said them. A line missing the quote comes back blank, because a verdict nobody can locate is worse than no verdict: it spends trust and returns nothing.

The design reason is stronger than disputes. A 2024 study of rubric-based scoring by language models found that asking a model directly for an overall quality score performed worse than simply predicting the average for every conversation: root mean squared error of 0.901 against 0.82 for the constant baseline, with a correlation to human judgment of 0.143. The same model on the same conversations reached 0.422 error and 0.350 correlation once the question was broken into a nine-part rubric, eight narrow dimensions plus the overall question, and calibrated to each individual human rater. That study scored information-seeking dialogue with a 2024-era model, not sales calls, which is why the thing to take from it is an instruction about how to build a scale and not a number about sales accuracy.

The instruction is specific. Do not ask for a verdict on the conversation. Ask twenty small questions with checkable answers, each of which a person could argue with, and assemble the number afterwards.

Apply that to whatever form you already use: every line opens, the composite does not. In Spellit Actions each scored line carries its quote as a button that drops into the transcript at that second and highlights it, while the number at the top opens nothing, because it is arithmetic over the lines and not an observation in its own right. If your dashboard lets somebody click a composite and expect an explanation, you have promised a drilldown with no evidence behind it.

One rule that costs nothing and prevents a whole class of argument. Where there is no data, show no colour. A call that was never scored shows a dash where a zero would go, and a rep with nothing to review this week gets a sentence saying so, in place of an empty chart. We shipped the opposite first: a block with no data for a particular client rendered as a row of zeros, and the screen announced a collapse where there had been nothing to measure.

An additive form passes a call that failed on the one thing that mattered

Which of these two forms passes a conversation where the rep let a wrong statement about the contract stand? The additive one does, comfortably, because four easy things were done well. Build the score by taking points away from a hundred instead, and one omission can decide the call. That is how a pipeline behaves anyway.

What the review form checksAdditive formSubtractive score
Opened with a clear agenda1 of 1nothing deducted
Covered the discovery list1 of 1nothing deducted
Presented the right part of the product1 of 1nothing deducted
Answered the question about commercial terms straight1 of 1nothing deducted
Let a wrong statement about the contract stand0 of 1deduct what you set for it
Result4 of 5, 80%, passesthe call is the problem

A default worth starting from is a quarter of the score for that last line, and the number belongs to you. In our implementation every weight comes from the client's own settings and not from our code. That is the only defensible arrangement: what is unforgivable in regulated advice is a rounding error in transactional sales.

Two shapes are worth marking absolute: the thing with legal consequences, and the thing that makes the rest of the conversation meaningless. Everything else is merely expensive. If every line is absolute, nothing is, and the form gets quietly ignored. Which two you pick is a decision about your own business, usually taken once and never revisited. Bring the two you picked and we will tell you what they cost you on real conversations.

What leaves the dashboard decides whether the review paid for itself

Feedback that names the person makes performance worse about as often as it makes it better. Feedback that names the sentence does not, and that is the whole difference between a score with somebody's name on it and an observation about a moment. The standard meta-analysis of feedback interventions, published in 1996 and still the reference, covers 607 effect sizes and 23,663 observations: an average improvement of d = .41, and more than a third of interventions making performance worse.

The authors' explanation is the part to act on. Feedback that turns attention toward the self degrades performance; feedback that keeps attention on the task improves it. A per-call score with a name and a number next to it is close to a worked example of the first kind.

So the thing to route is one line of the transcript, the sentence as it was said, the sentence that would have worked, and the second in the recording where it happened. A default worth starting from: one such line per rep per week, and none at all for a rep whose last three calls already show the behaviour. A second line in the same week does not double anything, it turns the first into background. Which line to pick, and what to do with it once it lands, is the weekly coaching loop that decides what actually gets sent.

Disputes belong in the design, and the appeal mechanics sit in that same coaching piece. The part specific to a review form is narrower: keep a list of which lines get challenged most, because a line that keeps getting disputed is more often a badly worded line than a run of difficult people.

The ways this goes wrong, all of which we shipped

A dashboard on one project reported that nearly every call contained a serious mistake. The real figure was a fraction of that, and nothing had crashed. Our service decided whether a call contained a serious mistake by looking for particular marker values, and different clients encoded the same fact in incompatible ways: for some an empty marker meant no mistake, for others a specific flag meant exactly that. Our list knew one convention and not the other. The number simply arrived, looked plausible in the way a bad number does, and sat on a screen somebody was about to make decisions from.

This is a different failure from the one described in the coaching piece, where the number was correct and useless. Here it was wrong.

The second is quieter. A rep's team membership was recorded one way in one place and another way in another, because a trainee can sit in two teams at once and only part of our system had been updated for that. Code written for the first way did not fail against the second. It returned a wrong roll-up, with no error anywhere.

Once a month, take the three numbers your dashboard shows largest and find the raw rows behind each one. A defect in call review does not crash. It arrives looking like a working answer, which is also the argument for scoring nothing without a quotable sentence behind it: a score you can open lands on something a person said at a particular second, and the person it describes can check it on the spot.

What to do next

Pick one line of your review form and check whether its score can be traced to words somebody said. If it cannot, that line is an opinion with a number attached, and it is the first one to rewrite. Work through the whole form that way before adding any coverage, because reviewing ten times as many calls against an untraceable form produces ten times as much of the wrong thing.

Then write the routing sentence: who gets what, in what form, on which day. One line per rep per week is a defensible starting point, and a rep already doing the thing gets nothing that week. Coverage without that sentence is a screen; coverage with it is a process.

Before you rebuild the form, measure the one you have. We will run your existing review form against a week of your recordings and hand back two lists: the lines that can be traced to a quote, and the lines that cannot. You keep both lists whatever you decide.
Key points
  • Whole categories of conversation never reach a hand-picked shortlist, and they are the same categories every week.
  • A sample sized for a team rate answers a question about the team. Managers ask questions about one person.
  • More than a third of feedback interventions in the research made performance worse, and the ones aimed at the person are the kind that does it.
  • In the one published test of this, asking a model for an overall conversation score did worse than predicting the average. Narrow, checkable questions worked.
  • Build the score by subtracting. A form that adds points rewards filling it in.
See the demo now. Then run it on yours.

A live demo of our products and a real conversation about the growth problem you need to solve.

Book a demo

FAQ

How many sales calls should a manager review each week?

Fewer than the number that gets scored, and chosen by something other than what reached the top of the queue. If every conversation is scored automatically, a manager opening three or four a week is reading the ones a score pointed at. How the shortlist was produced matters far more than its size.

Can AI score a sales call accurately?

On narrow checkable questions, yes. On holistic verdicts, treat it as unproven. The closest published evidence is not about sales: a 2024 study of information-seeking dialogue found that asking a model for an overall quality score did worse than predicting the average, while the same model became usable once the question was split into a nine-part rubric and calibrated to each human rater.

What should a sales call review form include?

For each line: what is being checked, in words a rep could disagree with; whether missing it fails the call outright or only costs points; and the rule that nothing is scored without a quotable sentence from the recording. Forms usually fail because their lines describe impressions, and impressions cannot be traced to anything anybody said.

Is reviewing every call overkill if a sample is statistically valid?

For a team-level rate, three hundred conversations pin it to about five points either side, and scoring everything narrows that to a fraction of a point. For a question about one person, the useful sample is that person's own month, about thirty conversations in B2B. Full coverage earns its keep on questions about individuals.

What makes a call score defensible when a rep disagrees with it?

It names what was checked, gives the verdict, quotes the sentence it rests on, and points at the second in the recording where that sentence occurs. A line with no locatable quote comes back blank. Then a human decides the challenge, and the system does as it is told.

How long should reviewing one call take?

About five minutes, once the call is scored and every line carries a timecode, because the manager opens the parts a score flagged. Listening end to end is still right for a conversation that will be discussed in detail. It should be a choice somebody makes, and not what happens by default because there is nowhere else to start.