How to review every sales call, and why coverage on its own changes nothing
A sales call review process at full coverage scores every conversation automatically, then routes a small number of specific observations to the people who are not already doing the thing. Coverage is the easy half and moves little by itself: a 2023 meta-analysis of workplace monitoring finds no sign that monitoring improves performance. In Spellit Actions the score starts at 100 and every failed line takes points off it.
Sales leaders listen to calls. They listen carefully, they take notes, and they hear things no review form will ever catch. That is the part most write-ups on this subject get wrong when they open with a line about managers who cannot be bothered.
The constraint was never effort. A week contains more conversations than one person can read, and the queue decides which ones get opened. A complaint pushes a call to the top. So does an escalation, a deal that visibly went sideways, a new starter whose first month somebody promised to watch. That ordering is the only one the week offers, and it is a perfectly sensible way to spend four hours. It is just not a sample. The arithmetic of that gap is in the coaching piece; this article is about what the conversations nobody opened have in common.
Whole categories of call never reach the shortlist
A queue built from complaints and escalations returns conversations that already announced themselves. The categories it skips are consistent, and each one holds a different kind of information. Opening more of the same calls does not close that gap.
| Call that rarely gets reviewed | Why it never surfaces | What it would have shown |
|---|---|---|
| A routine win | Nothing went wrong, so nobody flagged it | What the team actually does when it works |
| A quiet loss with no complaint attached | No escalation, no ticket, no noise | Where deals leak without anybody objecting |
| A call that ended in four minutes | Too short to look interesting | Whether qualification is cutting off the right people |
| A strong rep having an ordinary day | That name is not the one anybody is watching | Whether a problem has started spreading |
| A call in the team's second language | Harder to review, so it waits | Whether the script survives translation |
Run the check on your own team before changing anything. Take the last twenty calls anyone reviewed and write, next to each, what put it in front of them. If the answers are mostly incidents and a handful of names, the process is producing an audit of a shortlist.
A statistically sized sample does fix the arithmetic, and it deserves a straight answer. The numbers below are ours, worked from the standard margin on a proportion; no survey was involved. Three hundred conversations pin a team-level rate to roughly five points either side. Scoring all thirty thousand of a contact center's monthly calls narrows that to about half a point. For the question "how often does this team do X", the other twenty-nine thousand seven hundred buy almost nothing.
The catch is which question the rate answers. A leader rarely needs the team average to a decimal place. They need to know whether this person, on these conversations, keeps doing one specific thing, and for that the useful sample is the whole of one rep's month: about thirty conversations in B2B. At that size the statistically adequate sample and the population are the same set, and the case for sampling quietly disappears. Bring us a week and we will tell you which rows of that table your own review has never opened.
Coverage on its own moves nothing
Reviewing every call does not by itself improve how anybody sells. A 2023 meta-analysis of electronic performance monitoring, pooling 94 independent samples and 23,461 people in total, reports no evidence that monitoring improves worker performance. Across the 41 samples that measured performance, covering 5,804 people, the pooled correlation is −.01 with an interval of [−.05, .04] sitting squarely across zero.
What monitoring does move is measurable and unhelpful. The same analysis puts the correlation with stress and strain at .16, interval [.13, .20], and with felt invasion of privacy at .28, interval [.13, .43] from only eight samples. Workers' attitudes at work drift negative overall, −.11 with an interval of [−.20, −.03]. Monitoring people in more ways at once was the one thing that significantly predicted more counterproductive behaviour, though on four samples and an interval that barely clears zero.
Two honest caveats, because this is the paper most likely to be quoted back at anyone selling this idea. Its performance samples are dominated by experiments, with field sales teams barely represented. And monitoring for developmental purposes, which is the case this article is about, is the one purpose the authors could not test against performance at all: they found two studies. What the paper does kill is the claim that watching everything makes people sell better, and it kills it thoroughly, because it separately analyses monitoring that is recorded and reviewed later and finds the same nothing there.
So before switching coverage on, write down who receives what, in what form, on which day. If that sentence does not exist yet, full coverage will produce a screen nobody opens, and the honest reading of the evidence above is that the screen will change no behaviour while raising everybody's stress.
A line with no quote behind it should come back blank
Every scored line needs four things attached: what was being checked, the verdict, the words the rep actually said, and where in the recording they said them. A line missing the quote comes back blank, because a verdict nobody can locate is worse than no verdict: it spends trust and returns nothing.
The design reason is stronger than disputes. A 2024 study of rubric-based scoring by language models found that asking a model directly for an overall quality score performed worse than simply predicting the average for every conversation: root mean squared error of 0.901 against 0.82 for the constant baseline, with a correlation to human judgment of 0.143. The same model on the same conversations reached 0.422 error and 0.350 correlation once the question was broken into a nine-part rubric, eight narrow dimensions plus the overall question, and calibrated to each individual human rater. That study scored information-seeking dialogue with a 2024-era model, not sales calls, which is why the thing to take from it is an instruction about how to build a scale and not a number about sales accuracy.
The instruction is specific. Do not ask for a verdict on the conversation. Ask twenty small questions with checkable answers, each of which a person could argue with, and assemble the number afterwards.
Apply that to whatever form you already use: every line opens, the composite does not. In Spellit Actions each scored line carries its quote as a button that drops into the transcript at that second and highlights it, while the number at the top opens nothing, because it is arithmetic over the lines and not an observation in its own right. If your dashboard lets somebody click a composite and expect an explanation, you have promised a drilldown with no evidence behind it.
One rule that costs nothing and prevents a whole class of argument. Where there is no data, show no colour. A call that was never scored shows a dash where a zero would go, and a rep with nothing to review this week gets a sentence saying so, in place of an empty chart. We shipped the opposite first: a block with no data for a particular client rendered as a row of zeros, and the screen announced a collapse where there had been nothing to measure.
An additive form passes a call that failed on the one thing that mattered
Which of these two forms passes a conversation where the rep let a wrong statement about the contract stand? The additive one does, comfortably, because four easy things were done well. Build the score by taking points away from a hundred instead, and one omission can decide the call. That is how a pipeline behaves anyway.
| What the review form checks | Additive form | Subtractive score |
|---|---|---|
| Opened with a clear agenda | 1 of 1 | nothing deducted |
| Covered the discovery list | 1 of 1 | nothing deducted |
| Presented the right part of the product | 1 of 1 | nothing deducted |
| Answered the question about commercial terms straight | 1 of 1 | nothing deducted |
| Let a wrong statement about the contract stand | 0 of 1 | deduct what you set for it |
| Result | 4 of 5, 80%, passes | the call is the problem |
A default worth starting from is a quarter of the score for that last line, and the number belongs to you. In our implementation every weight comes from the client's own settings and not from our code. That is the only defensible arrangement: what is unforgivable in regulated advice is a rounding error in transactional sales.
Two shapes are worth marking absolute: the thing with legal consequences, and the thing that makes the rest of the conversation meaningless. Everything else is merely expensive. If every line is absolute, nothing is, and the form gets quietly ignored. Which two you pick is a decision about your own business, usually taken once and never revisited. Bring the two you picked and we will tell you what they cost you on real conversations.
What leaves the dashboard decides whether the review paid for itself
Feedback that names the person makes performance worse about as often as it makes it better. Feedback that names the sentence does not, and that is the whole difference between a score with somebody's name on it and an observation about a moment. The standard meta-analysis of feedback interventions, published in 1996 and still the reference, covers 607 effect sizes and 23,663 observations: an average improvement of d = .41, and more than a third of interventions making performance worse.
The authors' explanation is the part to act on. Feedback that turns attention toward the self degrades performance; feedback that keeps attention on the task improves it. A per-call score with a name and a number next to it is close to a worked example of the first kind.
So the thing to route is one line of the transcript, the sentence as it was said, the sentence that would have worked, and the second in the recording where it happened. A default worth starting from: one such line per rep per week, and none at all for a rep whose last three calls already show the behaviour. A second line in the same week does not double anything, it turns the first into background. Which line to pick, and what to do with it once it lands, is the weekly coaching loop that decides what actually gets sent.
Disputes belong in the design, and the appeal mechanics sit in that same coaching piece. The part specific to a review form is narrower: keep a list of which lines get challenged most, because a line that keeps getting disputed is more often a badly worded line than a run of difficult people.
The ways this goes wrong, all of which we shipped
A dashboard on one project reported that nearly every call contained a serious mistake. The real figure was a fraction of that, and nothing had crashed. Our service decided whether a call contained a serious mistake by looking for particular marker values, and different clients encoded the same fact in incompatible ways: for some an empty marker meant no mistake, for others a specific flag meant exactly that. Our list knew one convention and not the other. The number simply arrived, looked plausible in the way a bad number does, and sat on a screen somebody was about to make decisions from.
This is a different failure from the one described in the coaching piece, where the number was correct and useless. Here it was wrong.
The second is quieter. A rep's team membership was recorded one way in one place and another way in another, because a trainee can sit in two teams at once and only part of our system had been updated for that. Code written for the first way did not fail against the second. It returned a wrong roll-up, with no error anywhere.
Once a month, take the three numbers your dashboard shows largest and find the raw rows behind each one. A defect in call review does not crash. It arrives looking like a working answer, which is also the argument for scoring nothing without a quotable sentence behind it: a score you can open lands on something a person said at a particular second, and the person it describes can check it on the spot.
What to do next
Pick one line of your review form and check whether its score can be traced to words somebody said. If it cannot, that line is an opinion with a number attached, and it is the first one to rewrite. Work through the whole form that way before adding any coverage, because reviewing ten times as many calls against an untraceable form produces ten times as much of the wrong thing.
Then write the routing sentence: who gets what, in what form, on which day. One line per rep per week is a defensible starting point, and a rep already doing the thing gets nothing that week. Coverage without that sentence is a screen; coverage with it is a process.
Before you rebuild the form, measure the one you have. We will run your existing review form against a week of your recordings and hand back two lists: the lines that can be traced to a quote, and the lines that cannot. You keep both lists whatever you decide.
- Whole categories of conversation never reach a hand-picked shortlist, and they are the same categories every week.
- A sample sized for a team rate answers a question about the team. Managers ask questions about one person.
- More than a third of feedback interventions in the research made performance worse, and the ones aimed at the person are the kind that does it.
- In the one published test of this, asking a model for an overall conversation score did worse than predicting the average. Narrow, checkable questions worked.
- Build the score by subtracting. A form that adds points rewards filling it in.
A live demo of our products and a real conversation about the growth problem you need to solve.
Book a demo