Not applicable in call scoring is a number of its own, and nobody reports it

Most guidance on call scoring gets the arithmetic right: a line the conversation never called for leaves the denominator. The guidance stops there, before what that does to a summary report. Across nine sales teams at six companies in September 2026, on 83,671 scorecard rows in Spellit, the share of yes-or-no lines left blank ran from 2.1% to 81.6%. Two percentages built that way are not comparable.

Aram BelinskyCOO, Spellit11 min read

The advice on this question is good, which is the first surprise. Open the pages that discuss what to do with a line the conversation never called for and most give the same correct answer: take it out of the denominator and do not penalize a rep for a step the call did not require. Follow that and you will score one call correctly. Follow it for a year across two teams and you will have built two numbers that look alike and measure different things.

An earlier piece here, on what contact centers measure on every call and what they sample, ended on a rule we still stand behind: before anyone compares a rate across two teams, somebody has to be able to say out loud what each one divides by. This is what happens once somebody can. You get two honest answers, they turn out to be different answers, and the comparison the slide was built for has already been lost by then.

The published answer is correct, and it stops at the first call

Correct exclusion is a rule about a single conversation: this line did not apply here, so it leaves both halves of the fraction. The rule holds there. It stops holding at the point of aggregation, because once the exclusion has fired on eight lines in ten for one team and on almost none for another, the two resulting percentages have different things underneath them.

The governing document for contact center operations does not close that gap, which matters because it is the document people reach for. The COPC standard for customer operations, release 7.0, written by an organization that also sells certification against it, runs to 93 pages. Section 2.7 asks that the monthly number of interactions monitored for each program be based on an understanding of the statistical implications of the sample size, and that those doing the monitoring be calibrated to ensure consistency. Read the whole text and the phrase "not applicable" appears zero times, "N/A" zero times, and the word "denominator" once.

The sampling question gets 93 pages. The question of what the sample is then scored against gets one word, used once. Every form built to satisfy the standard inherits that asymmetry: an operation can certify with a form its conversations answer almost in full, and so can an operation whose form they barely touch, and the audit is not looking at the difference between them. Which of the two an operation is cannot be settled from its reports either, because a report carries the percentage, and the conversations it came from stay behind it. The conversations behind a month of them do settle it.

The share that goes unanswered is not a property of the form's length

Over September 2026, across nine sales teams at six companies and 83,671 scorecard rows, the share of yes-or-no lines left blank ran from 2.1% to 81.6%: 19,402 blanks against 64,269 scored. The team with the longest form, at 128 lines, sits near the bottom of that range.

TeamLines scoredLines blankShare blankYes/no lines on the form
Team 1 (see caveat below)2711,20481.6%14
Team 27,6504,11034.9%24
Team 321,9899,88231.0%29
Team 410,7852,36218.0%51
Team 55,25359710.2%25
Team 6746546.8%128
Team 714,9821,0576.6%40
Team 81,9041216.0%25
Team 9689152.1%43

Method, so the figures can be argued with: calendar month of September 2026, pulled on October 4, 2026; only lines carrying a yes-or-no scale, meaning the line has both 1 and 0 values and no text values anywhere in the month; share equals blanks divided by blanks plus scored; teams under 500 scorecard rows for the month dropped. The obvious objection is that a blank might mean no record was written at all, so we checked: a scorecard row is written for every call, and in eight of the nine teams the number of rows per call is a whole constant, 37, 40, 44, 76 and so on. In those teams a blank is a blank score, not a missing row.

Team 1 carries a caveat and it goes wherever the number goes. Its form changed partway through September, and only fourteen yes-or-no lines survived the whole month, so its 81.6% rests on 1,475 scorecard rows, a slice of the team's month. The figure is real and it is narrower than the others in the table.

Read down the last two columns and no relationship appears. Team 6 has 128 yes-or-no lines and leaves 6.8% of them blank. Team 1 has fourteen and leaves 81.6%. Team 4 has 51 and leaves 18%; Team 9 has 43 and leaves 2.1%. The blank share tracks fit between the form and the work, and neither shortening the form nor lengthening it improves that fit.

A score built on a fifth of the form is a different kind of number

Every form has a nominal size and an effective one. Nominal is how many lines are printed on it. Effective is how many a given conversation actually answers, and the percentage is built from the effective one. In the September data the effective size ranges from roughly 2.6 lines per card to roughly 119.

The arithmetic is plain. Team 1 answers 18.4% of fourteen lines, which is about 2.6 answers per card. Team 6 answers 93.2% of 128, which is about 119. Both produce a percentage, both percentages are correctly computed, and one of them is assembled from forty-five times more material than the other.

That difference shows up as movement. A rate assembled from two or three answers per card swings further from one month to the next than a rate assembled from a hundred and nineteen, for no reason except that there is less of it underneath. Put the two on the same trend chart and the short one looks erratic. Reading that wobble as instability is how a team gets asked to explain variation that nothing caused.

A default to start from, and this is our recommendation and not a measurement: print the answered count beside every rate, and treat a line left blank on more than half of a team's cards as a line written for different conversations, and move it.

None of that needs us, and it is worth saying so plainly because the rest of this piece reads like it might. Export last month's cards, group the rows by line, count blanks against scored, and you have the whole thing: a pivot table and an afternoon. A companion piece runs exactly that count down to one line, the next step question, where a team's blanks turn a 1.1% rate into 0.3%, depending on nothing but how they were counted.

Two blank cells, two different reasons

A blank cell means one of two things and they are opposites. Either the line did not apply to this conversation, which is a fact about the conversation, or nobody could tell either way, which is a fact about the evidence. The first is a clean exclusion. The second is an unanswered question wearing the same costume.

An earlier piece set out the rule we use on the second kind: a scored line with no locatable quote behind it comes back blank, because a verdict nobody can point to in the recording spends trust and returns nothing. That rule is right and we keep it. On its own it cannot hold the two kinds of blank apart once they reach a summary. Separating them after the fact means going back to the audio line by line, which is slow enough by hand that nobody does it twice, and it is the first pass we run over a set of scored calls.

Healthcare measurement at least gives the two kinds of removal separate names. The CMS guide for reading electronic clinical quality measures, version 10.0, May 2024 defines a denominator exclusion as removal from the denominator before the numerator is calculated, on the grounds that the numerator event is not applicable, and a denominator exception as removal after the numerator has been counted. No reporting requirement is attached to exclusions at all. For exceptions it is conditional: when exception cases are removed, the measured entity may still be required to report how many patients or episodes carried a valid one. That is a thinner provision than it first sounds, and it is still two named objects and one conditional count more than call scoring has.

Why the cell is blankWhat it is a fact aboutWhere it belongs in the arithmetic
The conversation never reached this stepthe conversationOut of the denominator, counted and printed separately
The recording or transcript does not settle itthe evidenceIn the denominator as unanswered, or the line is withdrawn and the withdrawal is reported

Store both as the same blank and a specific failure goes invisible. Audio quality drops on one team, more lines come back unverifiable, those lines leave the denominator alongside the genuinely inapplicable ones, and the completion percentage rises. The score improves because the evidence got worse. Nobody in that chain did anything wrong, and the number on the slide is now moving in the wrong direction for a reason nobody in the room can name.

What we got wrong while counting this

The first version of the summary table behind this export counted a blank as a zero. Nothing broke; the ordering just looked wrong, and that is how we caught it. No customer saw a row of zeros: this was our own working table of these nine teams, and the mistake changed the order the teams sat in. Run the same September data both ways: the team at 2.1% blank barely moves between the two versions, and the team at 81.6% drops from the middle of the table to the bottom of it. Nothing about anyone's work differed between the two tables. We had produced a ranking of how much of each form the conversations happen to touch, formatted as a ranking of how well each team sells.

The second mistake cost a day and is less flattering. The first count pooled every kind of line together: yes-or-no lines, lines whose answer is a set of labels, and lines whose answer is free text. On a yes-or-no line a blank means the question did not arise. On a free-text line a blank mostly means the conversation produced nothing worth writing down, which is a different event and a far more common one. Pooled into one figure the two inflated the result for exactly the teams whose forms lean on text. We split lines by the shape of their answer and recounted, and every figure in this piece is from the recount, binary lines only. The number we almost published was higher, and once published it would have been quoted.

One more thing belongs in this section. The piece passes judgment on a whole field for not printing a number, and the number it prints came out of our own system, where it is not printed either. It exists here because somebody wrote a query by hand across a month of scored cards to produce it. That is the honest status of the figure: a measurement taken on purpose for this article, not a line anybody was already reading. The same applies to the discipline it argues for, which is why the answered count belongs beside the lines it describes, on a scored conversation.

What to do next

Take one month of scored cards for two teams you have been comparing. For each team count two things: how many yes-or-no lines were answered and how many were left blank. Write the blank share next to that team's quality percentage, on the same line of the same slide. If the two shares differ by more than a few points, the comparison those percentages were supporting has to be rebuilt line by line before it means anything.

Then take the line with the highest blank share in either team and ask what it is doing on that form. A question answered on a fifth of the cards measures how often that team's conversations happen to resemble somebody else's. A form written for outbound selling, pointed at an inbound queue, will be blank most of the way down, and no tolerance setting repairs that.

A form that fits one team's conversations is the wrong instrument for another's, and the blank share is where that shows up first. Before anybody's target moves, find out which lines your conversations actually reach and which were written for work somebody else does. A month of calls names the lines yours never reach.
Key points
  • Removing a line the conversation never called for from the denominator is correct, and most published guidance says so.
  • Doing it correctly is what makes two teams' percentages incomparable: each one ends up computed over a different subset of the form.
  • Across nine sales teams at six companies in September 2026, the share of yes-or-no lines left blank ran from 2.1% to 81.6%, on 83,671 scorecard rows.
  • The length of the form does not predict that share. The 128 line form in the set is answered on 93% of its lines; the 14 line form on 18%.
  • A blank cell has two possible causes, and a form that stores them identically lets falling recording quality raise the score.
See the demo now. Then run it on yours.

A live demo of our products and a real conversation about the growth problem you need to solve.

Book a demo

FAQ

Why does excluding inapplicable lines correctly make two teams harder to compare?

Because the exclusion is applied per conversation, each team ends up with a percentage computed over whatever subset of the form its own conversations activated. One team's 85% can rest on nearly every line of a 128 line form; another's can rest on two or three answers out of fourteen. Both are correctly computed, they are not the same quantity, and the headline numbers do not carry the difference.

How should the share of inapplicable lines be reported?

As a count next to the score, per team and per line, for the same period. A quality score of 85% with 4% of lines blank and one with 60% blank are different measurements printing the same digits. Healthcare measurement is a little further along: it gives the two kinds of removal separate names, denominator exclusion and denominator exception, and for exceptions a reporting entity may still be required to state how many cases carried one. Call scoring has neither the names nor the count.

Can two teams' quality percentages be compared at all?

Line by line, and only for lines both teams answer often enough to carry a rate. The headline percentages describe different fractions of the form, so on their own they compare nothing. Comparing a single line across two teams, with the answered count printed beside each, is the version that holds, and it is also the version a coaching conversation can act on.

What share of blank lines is too high?

No published threshold exists, so here is ours: treat a line left blank on more than half of a team's cards as a line written for different conversations, and move it instead of tolerating it. That is a rule of thumb taken from the spread we measured, nine teams running from 2.1% to 81.6% blank in September 2026, not a threshold anyone has validated. At the top of that range the form is pointed at the wrong work.

How is a line that does not apply different from one with no evidence either way?

The first is a fact about the conversation: the step never came up. The second is a fact about the recording: the moment may well have happened and nothing in the material settles it. Only the first is a legitimate exclusion. Treat the second the same way and a drop in recording quality quietly raises the score, because unverifiable lines then leave the denominator along with the inapplicable ones.

Does the 85% industry benchmark mean anything for my team?

Not as a target. The widely quoted figure comes from [a published list of customer service measures](https://www.sqmgroup.com/resources/library/blog/7-essential-customer-service-metrics-and-how-you-measure-them), which gives a quality score benchmark of 85%, with 90% to 99% described as good, and says nothing about how many lines the score was computed over or what happens to a line that does not apply. Your 85% computed over a fifth of your form and a published 85% computed over all of someone else's are not the same measurement. Use your own trend instead.