Next step statistics across 6,819 sales calls, and what the spread is describing

Across 6,819 scored sales calls at seven companies in the 90 days to October 4, 2026, Spellit recorded an agreed next step on 28.3% of them, with a team median of 34.9% and individual teams running from 1.1% to 80.9%. That spread is not a league table of selling ability. It sorts by the kind of conversation each team holds: first consultations near the top, an inbound contact center queue at the bottom.

Aram BelinskyCOO, Spellit11 min read

The usual move with a number like this is to find one percentage somewhere, call it the benchmark, and set it as a target by Monday. Our own figures say why that goes wrong, and the problem is not that the number is hard to find: there are thirteen of them, every one is correct, the highest is more than seventy times the lowest, and the shape they make is worth more than any one of them.

Thirteen teams sort by type of conversation, not by company

Sorted by share, these teams do not line up by company size or by industry. They line up by how early in a purchase the conversation sits. Teams whose job is a first consultation occupy the top of the table. Teams working an inbound queue or a call that follows a written proposal occupy the bottom. The line being counted is named differently in every company's form, so this counts the same question and not the same label: was a next step fixed on this call, scored yes or no, with blanks held out of both halves of the fraction.

TeamShare with a next step recordedCalls scored
Company A, first consultations80.9%314
Company A, follow-up calls65.8%883
Company B, three sales teams together61.7%133
Company B, follow-up calls46.5%286
Company C, first product line34.9%436
Company D, international sales22.5%138
Company E, before the proposal16.2%1,710
Company E, after the proposal15.3%2,486
Company C, second product line14.4%188
Company G, inbound contact center1.1%175

Three cautions before anyone copies a row out of it. One team of 70 calls is held out of the table and three small teams at one company appear as a single row of 133, because a slice under a hundred calls starts to describe one identifiable customer rather than a pattern; all of them stay inside the 6,819, the median and the pooled figure. The pooled 28.3% is dominated by company E, whose two slices are 4,196 of the 6,819 scored calls and both sit near 15%, which drags it well below the median team. And these seven companies are our customers, which is a selection effect we cannot correct for: a company that bought a tool for reading calls is more likely than average to have a line about the next step on its form in the first place, and nobody outside can check a row of this table against anything. Building the same table for your teams needs none of us: one filter, one count, an afternoon. The part worth handing over is the second opinion on which rows are comparable at all.

Inside one company, two teams sit twenty points apart on the same form

The clean test of whether a next step rate measures selling ability is to compare two teams inside one company, where the form, the window and most of the surrounding conditions are the same. There are four such pairs here. Three show gaps wide enough to change a decision made on them, and the fourth is interesting for the opposite reason.

Company C sells two different things through two teams, at 34.9% and 14.4%, a gap of twenty points inside one business. Company A runs first consultations at 80.9% and follow-up calls at 65.8%. Company B's three sales teams come to 61.7% together while its follow-up team sits at 46.5%. The same company, the same form, the same window, and two numbers that would get two different conversations in a review. For teams like company A's, where the whole engagement is the consultation itself, the top of the table is the normal reading rather than an achievement.

The fourth pair is the counter-example sitting inside our own table, and it is not a small one. Company E's two slices are 4,196 calls, nearly two thirds of everything here, and before and after a written proposal the share moves by nine tenths of a point. Whatever holds that team near one call in six, the stage of the deal is not it. Type of conversation is not a single dial that explains every row either.

Nothing in this data measures quality: there is no spread between reps inside a team, no conversion, no deal outcome, and no way to get any of them from a scorecard. So the claim stays narrow. None of this says selling ability leaves the number alone. It says the ranking cannot be read as a ranking of skill, and that is the only thing the table supports.

An unscored call is not a failed call, and one team is three quarters unscored

The lowest team in the table, an inbound contact center queue at 1.1%, carries 581 calls with no score on the next step line at all, against 175 calls scored either way. Blank does not mean somebody failed to fix a next step. It means the question does not apply to the call, which for an inbound queue is most of the time.

TeamScored yes or noLeft blank
Company G, inbound contact center175581
Company E, before the proposal1,710597
Company E, after the proposal2,486491
The other ten teams combined2,44825

Of the 756 calls where that line exists at all, 581 are blank, which is 76.9%. Exclude them and the team reads 1.1%. Count them as zeros, which is exactly what happens when a report insists on one number per team, and it reads 0.3%. Four times worse, from a decision nobody writes down. Company E carries the same risk at a smaller scale, at 25.9% and 16.5% blank, and the other ten teams carry 25 blanks between them and are unaffected either way.

That blanks mean "not applicable" is our reading, not a label the data carries. A separate count of 83,671 scorecard rows from September 2026 found that eight of nine teams write a whole, constant number of rows per call, so the row exists and the value in it is empty rather than missing. What that does not rule out is a reviewer skipping a line, or a form that changed mid window. Neither explanation stretches to three quarters of one team's calls, and for the teams at 25.9% and 16.5% it is a live possibility we cannot close. What a not applicable score does to every other line on a form is a piece of its own, and so is how contact center quality programs pick and score calls at all.

A mostly blank slice therefore gets published on its own, or not at all, with the blank count standing next to the share, and it does not get cleaned up before reporting. A target set on the share without the blank count beside it is cheap to announce and slow to unwind, because by the time it reads wrong the team has been working to it for a quarter. Half an hour settles which of your teams the question was written for.

A filter matched two scorecard lines with almost the same name and tripled a number

Pulling the contact center figure for this article, I got 3.3% on the first attempt and 1.1% on the second, and the difference was not in the data. I had filtered the export by the number of the checklist item rather than by its text, and that company's form carries two lines whose names are near duplicates of each other. The filter took both, and two populations got added together.

Nothing in the data prevents that. Each company writes this question in its own words, because the scoring runs against the checklist a company already uses instead of against a list of ours, and a position number is not an identifier that survives being copied between forms. A near duplicate line is an ordinary thing to find in a form that several people have edited over two years.

The two versions are further apart than they look. At 3.3% the range across these teams reads as a factor of twenty-four; at 1.1% it reads as a factor of more than seventy, and the opening sentence of this article moves with it. The repair is dull and permanent: match on the text of the line, not on its number, and print the names of everything the filter matched before printing any share.

It cannot be fixed further upstream. The names belong to each company's form and are theirs to edit, so anything keyed to a position goes wrong again the next time somebody reorders a list. A number pulled once for an article gets looked at. The same query scheduled weekly against a form other people keep editing gets looked at once, at the start, and then trusted for a year. That is where a position number does its damage.

A 412% uplift has nowhere to sit when the base rate is 35%

The most quoted figure on this subject says top performers are 412% more likely to have a next step or meeting defined. It comes from the 2024 B2B Sales Benchmark Report by Ebsta and Pavilion, built on 4.2 million opportunities and more than a million hours of conversations from 530 companies representing over $54 billion in revenue, covering 2023. It counts a field on an opportunity record rather than a sentence spoken on a call, which already makes it a different population from the one in the table here. And read as a relative uplift it needs a base rate, which the report does not print.

Work the arithmetic and the gap shows. A 412% uplift reads as 5.12 times as likely, so against a base of 10% the top group sits at 51%, and against a base of 19.5% it sits at 100% with nowhere above to go. Eight of our thirteen teams are above that line and the top one is at 80.9%. The report does not say whether the figure is a ratio of probabilities or a ratio of odds, and that matters: as odds the ceiling disappears, and a base of 35% lands at 73.4% and not at an impossibility. The objection holds under the first reading and dissolves under the second, which is itself the problem with printing the number bare.

Base rate for the comparison groupIf 412% is a ratio of probabilitiesIf it is a ratio of odds
5%25.6%21.2%
10%51.2%36.3%
15%76.8%47.5%
19.5%99.8%55.4%
35%, our median team179%, arithmetically impossible73.4%

None of that is a complaint about the sample. Four point two million opportunities is a serious amount of data, and that is the objection: a set that size is one of the very few that could publish a distribution, and it prints one average per behavior instead. The shape is what 530 companies of data could have contributed, and the shape is what got averaged away. A third set of figures, circulating from a conversation analytics vendor, pins one close rate drop to three different samples depending on which retelling you read, and tracing those numbers is a job already done next door.

A buyer survey is the other source cited around this subject, and it answers a different question. RAIN Group's Center for Sales Research surveyed 488 buyers responsible for more than $4.2 billion in purchases across more than 25 industries, and 489 sellers who prospect, and reports that buyers say 58% of their sales meetings are not valuable to them. The document carries no date on its cover; the research background page inside puts the field work in June and July 2017, nine years back, and the file was produced in February 2018. It is a recollection gathered by online panel rather than a count taken off recordings, so it cannot be stacked against the 28.3% here. RAIN sells prospecting training, and the document ends on a description of the course.

What to do next

Take your own last ninety days and resist computing one number. Split the calls by type of conversation first: first contact, follow-up, after a written proposal, inbound. Count each slice separately and print the number of calls behind every share right next to it. Any slice under a hundred calls is an anecdote with a percent sign on it and should be labeled as one.

Then print a second column beside the first: how many calls in each slice carry no score on this line at all. Where that column is large you have found the slice the question was never written for, and any average that swallows it is wrong in a direction you cannot predict. Before any of it, print the names of the lines your export matched. Two of ours were near twins, and finding that took one command.

Two columns, side by side, team by team. The share and the blank count belong next to each other, because either one alone will mislead whoever reads it. What a table of your own cannot supply is somebody else's ten rows to sit it against.
Key points
  • Thirteen teams, seven companies, 6,819 scored calls in one 90 day window. The range is 1.1% to 80.9%, the median team 34.9%, the pooled figure 28.3%.
  • Gaps inside a single company reach twenty points, which is wider than an explanation pitched at the level of company culture can carry.
  • The lowest team is low mostly because the question does not apply to its work: 581 of its calls carry no score on that line at all, against 175 that do.
  • Counting those blanks as zeros, which is what a one number per team report does, turns 1.1% into 0.3% and makes the team look four times worse than the data says.
  • The famous "412% more likely" needs a base rate its report does not print. Read as a ratio of probabilities it has room to exist only below about 19.5%, and eight of our thirteen teams sit above that.
See the demo now. Then run it on yours.

A live demo of our products and a real conversation about the growth problem you need to solve.

Book a demo

FAQ

What share of sales calls end with an agreed next step?

There is no single share, and the useful answer is a distribution. Across 6,819 scored calls at seven companies over 90 days, the pooled figure was 28.3%, the median team was 34.9%, and individual teams ran from 1.1% to 80.9%. Which end a team sits at tracks the type of conversation it runs more closely than anything else that can be read off a scorecard.

Why do two teams in the same company get different next step rates?

Because they are having different conversations. In one company here, one product line sat at 34.9% and the other at 14.4%. In another, first consultations recorded a next step on 80.9% of calls and follow-up calls on 65.8%. The form and the window were the same in both cases. Everything else, including who manages each team, is outside what a scorecard records.

Should a call with no score on this line count as a failure?

No, and counting it that way is the most common way these figures go wrong. A blank means the question did not apply to that call. One inbound contact center team here had 581 blank calls against 175 scored ones. Excluding the blanks it reads 1.1%; counting them as zeros it reads 0.3%, a difference produced entirely by a reporting choice.

What should I put in place of a benchmark?

Three things on the same row: the share, the number of calls it stands on, and the type of conversation the slice contains. Drop any slice under a hundred calls or label it as an anecdote. Add the blank count as a fourth column. A row carrying all four can be compared with next quarter's row; a bare percentage cannot be compared with anything, including itself.

Why is an inbound contact center team so low on this measure?

Mostly because a next step is not what most of those calls are for. In the slice here, 581 of the 756 calls where the line exists carried no score on it at all, which says the question was written for a different kind of conversation. The 1.1% is accurate as a count and close to meaningless as a performance figure.

How do I compare next step rates across companies with different forms?

Carefully, and by the text of the question rather than its position. Each company names this line in its own words, and a form can hold two lines with near identical names. Match on wording, print the names your export matched before you print any share, and state the call count and the window next to every figure you compare.