Three numbers hide behind the accuracy of a meeting note, and they do not predict each other

Asking a vendor how accurate their meeting notes are produces a number about word recognition, which is the one question you did not need answered. Speaker labelling fails on its own schedule, and the summary turns out to be almost indifferent to how badly the transcript went: across 1,760 measurements, the correlation between transcription error and summary quality topped out at around minus 0.5 and fell to minus 0.15 on some measures. Spellit Notes carries all of these.

Oleg KulakovCEO, Spellit8 min read

Splitting accuracy into stages is no longer the insight it was a year ago. The better write-ups do it, and so does the box above the search results. What none of them say is that the three numbers are largely independent of each other, which is the part that decides what you do about it.

Word error on a headset and word error in a room are two different numbers

The difference between a headset and a laptop on the table shows up on every model anyone has measured. The original Whisper paper, published by OpenAI in December 2022, reported word error of 16.9% on the headset-microphone version of the AMI meeting corpus and 36.4% on the same meetings recorded through a single distant microphone. The older model it was compared against showed the same shape: 37.0% against 67.6%.

The easy half has moved since. On the Open ASR Leaderboard, updated at the end of September 2026, the best systems sit at around 6% word error on the cleaned headset version of AMI. The hard half has not moved in the same way, and it is measured on a different corpus, so the two numbers do not subtract. The NOTSOFAR-1 challenge, run by Microsoft on 315 real meetings in 30 rooms with four to eight participants, saw the winning single-channel system reach 22.2% error. The same team's best system in the multi-channel track, using a dedicated microphone array rather than one device on the table, reached 10.8%.

That last comparison is the useful one, because the team and the data were held constant and the recording condition is what changed. The practical instruction follows directly and costs nothing: join the call through the platform rather than putting a laptop in the middle of a table. Everything a meeting record is later used for inherits that choice. Everything downstream inherits whichever of those two numbers you chose, and your own recordings carry the answer to which route each one took.

Who said what fails on its own schedule

Speaker labelling fails on its own, and fixing word recognition does not fix it. The NOTSOFAR-1 results report both: error with speaker attribution required, and error with speakers ignored. The winning single-channel system scored 22.2% and 17.7%. One other entrant scored 40.1% and 27.2%, meaning thirteen of its forty points of error were purely the question of who was talking.

Measured on its own, a 2025 benchmark from ETH Zurich ran five diarization systems over 196.6 hours of audio in five languages, counting overlapping speech rather than excluding it. The best system, a commercial product, reached 11.2% diarization error; the best open-source alternative reached 13.3%. The authors traced the dominant cause to missed speech followed by speaker confusion, worsening as the number of participants rose.

This matters because of what gets attached to a name. An action item assigned to the wrong person is not a transcription error that a reader will catch; it is a correctly transcribed sentence with the wrong label, and it travels straight into whatever follows the meeting. The authors of Whisper noticed the related failure in their own model and wrote it down: it had a tendency to transcribe plausible but almost always incorrect guesses for the names of speakers.

The summary barely notices how bad the transcript was

Here is the measurement that decides what an accuracy figure is worth: somebody generated summaries from transcripts of every quality and scored them.

A 2025 review of the CHiME-7 and CHiME-8 challenges took the NOTSOFAR-1 evaluation set and generated summaries from each team's best transcript, eight summaries per transcript with different seeds, then scored those summaries. That is 1,760 system-and-session pairs. The systems ranged from about 11% word error to about 70%.

The summaries did not track that range. Correlation between transcription error and summary quality came out at minus 0.54 on the strongest measure and minus 0.15 on others. In the authors' words, systems where roughly one word in two was correctly recognized and attributed produced summaries roughly on par with systems that got about 80% right. Their own wording: meeting summarization may not be a good proxy for transcription quality, because of how well it survives transcription errors.

The detail that should make you uncomfortable is what happened to the deliberately broken controls. The researchers built fake systems by randomly deleting half the reference words, and by randomly reassigning every speaker. Summary quality dropped on those, as you would hope. Fluency did not. The summaries read smoothly in every condition, including the ones built from wreckage.

What you can check by reading the noteWhat it tells you
Does it read wellNothing. Fluency held up on transcripts with half the words deleted
Does it sound confidentNothing
Does the action item appear under the right nameSomething, and it is the number that fails independently
Does the timestamp lead to the moment it claimsSomething you can verify yourself, and the only one on this list you can

What the measurements support is narrower than that. Fluency carries no information and confidence carries none either, so the thing left to check is whether a claim can be traced back to a moment in the recording. Any tool that lets you do that passes; the figure on a website does not tell you whether it does. The same applies to a chatbot's read of a transcript. Which lines in a note of yours would survive that check?

What goes wrong is omission, not invention

A 2026 benchmark took 5,856 key facts across the meeting portion of 1,800 conversations, extracted by a model and adjudicated by four of the paper's own authors wherever the model rejected one, then counted how many survived into the summary.

It scored summaries on both faithfulness and completeness. Faithfulness was high: between 92.6% and 98.2% of statements in meeting summaries were supported by the conversation. Completeness was not. One widely used model captured 36.5% of a meeting's key facts and another reached 42.6%, while the weakest fell below 20%.

Its meeting transcripts are a mix: two of the three sources are human-prepared and one is speech recognition output from city council recordings. Its facts were verified by a model checked against graduate annotators on a sample rather than by annotators throughout. So this describes the summarizing step on inputs that are cleaner than yours, and it is the generous end of the range.

The consequence is a different reading habit. Checking a note for errors is checking for the rare failure. Checking it against what you remember happening is checking for the common one, which is the same reason a review process has to decide what it is sampling for. And nobody has measured how often a summary attributes something to a participant who did not say it when the input is real speech recognition output rather than a clean transcript. We looked, the gap is real, and one competitor says the same thing in its own write-up.

We hold a confidence number and do not show it

We hold a number we do not show you, and we have not resolved the argument about it.

When the system labels a line of a transcript with a speaker, it records how confident that label is. That number sits in the data and never reaches the screen. A reader sees a name next to a sentence, in the same typeface and with the same authority, whether the system was certain or was guessing between two people who sound alike.

We have not shipped it, for a reason that has not improved with time. Putting a confidence figure next to a name invites the opposite of caution: readers would stop checking the high numbers. We would be asking people to trust a threshold, and we have not measured what that threshold means on a real meeting with four participants and crosstalk, which is exactly the situation where it matters. So we have a number that would help, a worry about how it would be read, and no measurement to settle it. Meanwhile a reader of our notes sees a guessed name in the same typeface as a certain one and has no way to tell. That is the cost, it is being paid now, and it is paid by them rather than by us. The habit that does work in the meantime is the cheap one: open the timestamp and listen.

What to do next

Take ten of your own recorded meetings, the ordinary ones rather than the clean ones, and check three things separately. For each, pick one action item from the note and open the moment it points to. Does the audio contain the commitment. Was it the named person who made it. And is there anything else in that part of the conversation that the note left out entirely.

Three tally marks per meeting, ten meetings, under an hour. At the end you have a word-level number, a speaker number and an omission number for your own rooms, your own microphones and your own people. No published figure gives you any of the three, ours included, and the second one is usually the surprise.

Ten of your own recordings answer this better than any figure we could quote, ours included. Pick the ordinary meetings rather than the clean ones, and for each note we will mark which lines hold up against the moment they point to, and what the conversation contained that the note does not. The ordinary meetings, not the clean ones.
Key points
  • Word error on a headset and word error on a laptop in a meeting room differ by roughly a factor of two, on every model measured.
  • Who said what is a separate failure. On one benchmark system, 13 of its 40 points of error were speaker confusion alone.
  • A summary built from a transcript with half the words randomly deleted still scores high on fluency. It reads fine.
  • Meeting summaries invent little and leave out a great deal: one widely used model kept 36.5% of a meeting's key facts.
  • So the vendor's percentage does not convert into how much you can trust the note, in either direction.
See the demo now. Then run it on yours.

A live demo of our products and a real conversation about the growth problem you need to solve.

Book a demo

FAQ

How accurate are AI meeting notes?

There is no single figure, because three different things are being measured. Word recognition runs around 6% error on a headset and around 22% on one microphone in a meeting room. Speaker labelling fails separately, at roughly 11% on the best systems. And summary quality turns out to correlate only weakly with either, so neither number tells you whether the note is trustworthy.

Why do vendors quote numbers above 90%?

Because transcription on clean audio is the easy measurement and the flattering one. Those figures are usually taken on a single speaker with a close microphone, which is not the condition your meetings are recorded in. They also do not predict the quality of the summary built on top, which is the thing you actually read.

Do AI notes make things up?

Less often than people assume. On a 2026 benchmark of meeting summaries, between 92.6% and 98.2% of statements were supported by the conversation. The frequent failure is the opposite: one widely used model captured 36.5% of the meeting's key facts. The note is more likely to be missing something than to be inventing it.

How do I check whether a meeting note is right?

Open the timestamp. Every other signal is unreliable: fluency held up in a controlled study even on transcripts with half the words randomly deleted, so reading well tells you nothing. A claim you can trace back to a moment in the recording is checkable, and a claim you cannot trace is not evidence regardless of how it reads.

Does a better microphone matter more than a better model?

Usually, yes. In one challenge, the same team's best entry scored 22.2% error with a single device in the room and 10.8% with a microphone array, on the same meetings. Joining the call through the meeting platform instead of recording the room is the version of this that costs nothing.

Can I rely on who the note says was speaking?

Treat it as the weakest of the three numbers. On one benchmarked system, thirteen of its forty points of total error were speaker confusion alone, and errors rise as more people join. For anything that assigns an action to a named person, the attribution is worth confirming against the recording rather than against the note.