Running your calls through a chatbot works, and here is where it stops working
Pasting a transcript into a chatbot and asking what went wrong works, and it answers the question once. Turning that into something a team runs weekly breaks in three places: the evidence the model offers, the stability of the answer between runs, and what happens when you feed it a quarter instead of a call. Spellit Actions was built against the volume problem, and it does not escape the other two either.
A note on what this piece is not. It is not an argument that models are unreliable narrators and humans are not. In a blind comparison of clinical summaries, a model invented things on 5% of samples and the medical experts it was compared against did so on 12%. The question is not whether the thing works. It is where the edge is.
The advice on this subject contains no measurements at all
Searching for how to analyze a sales call with a chatbot returns almost exclusively vendors. Of the twelve distinct pages we read, eleven belong to a company selling transcription, meeting notes or call analysis. This page is the twelfth, and it is written by one of those companies. So the fact worth arguing about is not who wrote what: it is that not one of the twelve reports a sample, a period or a method for any claim it makes. The ones below do, and that is the only basis on which this piece asks to be read differently.
What they do contain is numbers that cannot be checked. One guide hands you a prompt that asks a model for an "effectiveness rating (1-5)" and a "probability estimate" from a single transcript, and says nothing about what either number would mean. Another, written entirely in the conditional, says the model "might provide a list of timestamps and quotes" and could offer a comparison against "successful pitches in your industry", which is data the model does not have.
The one page in the set that names failures rather than possibilities is blunt about them: a chatbot cannot join the call, will invent a qualification score from a partial transcript, and does not store the score anywhere. The same silence covers how good the transcript was in the first place. That company sells an alternative, so read it accordingly. It is still the only page in the set that describes what happens rather than what might. We can show you which of your own calls carry a quote and which carry a guess.
A model asked for proof will produce proof
The first thing people ask for, correctly, is evidence: do not just tell me the rep mishandled the objection, show me where. Models comply, and the compliance is the problem.
A September 2026 study from Johns Hopkins measured this precisely, in the setup that matters here: the documents are handed to the model, retrieval is frozen, and only the generator changes. Across twelve models and 222 questions, every claim was traced through a funnel. Does it carry a citation? Does the quoted string actually appear in the cited section, character for character? And does the quote support the claim?
The results split by model weight in a way worth knowing. A lightweight model returned quotes that could not be found in the source at all almost a third of the time. A frontier model in the set had the opposite profile and the more dangerous one: 98% of its claims arrived with a verified, genuine quote attached, and only 37.1% were judged fully supported by it. The citation is real. It does not prove the thing it is sitting under.
This is a preprint, the questions are synthetic, and support was judged by a model checked against two human annotators rather than by the annotators alone. The direction holds elsewhere: a Stanford and Yale team studying legal research tools coined the word misgrounded for the same pattern, where key propositions are cited and the source does not support them.
The authors then went back through the same sections looking for a different quote that would have supported the same claim. They usually found one. Support rose from 37.1% to 85.1%. The proof was in the document. One pass was bad at picking it out.
The same question, asked twice, gets two answers
Repeatability is the second place this breaks, and it breaks lower than people expect. A 2025 study across five models and eight tasks ran each condition ten times at temperature zero, top-p one, with the seed fixed. These are the settings people reach for when they want determinism.
On a professional accounting task, one model's best run scored 89.0% and its worst 57.8%, with a median of 74.5%. The share of questions where all ten runs returned a character-identical answer was 4.6%. On a college mathematics task, that share was zero.
Those are multiple-choice tasks where a right answer exists, which is the easy case, and nobody has published the equivalent for an open read of a transcript. If a task with a single correct answer cannot be reproduced at temperature zero, a task with no correct answer will not do better.
For a one-off question this does not matter. You read the answer, you use your judgment, you move on. It matters the moment somebody asks whether objection handling improved since last month, because the comparison needs the two measurements to have been taken the same way. This is the difference between an answer and a measurement, and it is the same line that separates a read of one call from a review process. Nothing about running it inside a product removes the underlying variance; what changes is that the question and its wording stop moving between runs. Bring the question you would want answered every month, and we will show you what has to be settled about it first.
Past a point, more context makes the answer worse
A study of how models use long inputs ran 2,655 questions with the answer-bearing document moved around the context and nothing else changed. Accuracy moved by more than twenty points depending on where it sat. Volume is the third place this breaks, and it surprises people who have been diligent about collecting recordings. In the twenty and thirty document settings, accuracy fell below what the same model scored with no documents supplied at all, which was 56.1%.
| What you do | What you get |
|---|---|
| Paste one transcript | A usable read of that call |
| Paste ten | Still usable, and now you are comparing by hand |
| Paste a quarter's worth | An answer that sounds the same, drawn from the parts the model happened to attend to, which is a different job from counting what customers keep saying |
| Paste a quarter's worth and ask a follow-up | A second answer, also confident, not necessarily consistent with the first |
A model's advertised context window is not the length it works at. On a 2025 benchmark built to remove word overlap between the question and the answer, one model advertising a 128,000-token window held its accuracy to about 8,000, and another advertising 200,000 held to about 4,000. That benchmark was built to be hard and its haystacks were novels, so read those as a floor. The number you need is your own and it costs twenty minutes: paste transcripts one at a time and stop when the answer stops getting better.
The proof our own system produced was decorative
We asked a model to justify its conclusions with a call number and a customer quote, which is exactly what the sensible version of this looks like. It returned beautiful, round, confident call numbers. In the incident that made us stop, the numbers were 12345 and 78901. Clicking through to the evidence crashed the screen, because the call did not exist.
The crash was the lucky part. A broken link announces itself. The review piece on this blog argues that a score needs a quote behind it; this is the other half of that lesson, which is that a quote can be present and be worth nothing. What we had actually shipped was a system where a conclusion arrived wearing proof, and the proof was decorative, and for a while nobody noticed because it looked right. The Stanford work on search engines found the same relationship from the other end: the more useful readers judged an answer, the worse its citations were.
The fix cost us something and we would make the same trade again. A quote now counts only if it is found in the transcript word for word. A call number is checked against the calls that exist. And when a conclusion cannot be grounded, we show the conclusion and hide the link rather than attaching the nearest plausible one. That last decision is the uncomfortable one: a reader sees an assertion with nothing to click, and asks why. The honest answer is that we could not prove it, and that is better than a number that dissolves when you press it.
What to do next
Run the same prompt on the same transcript tomorrow morning, before you check anything else. Put the two answers side by side and mark every place they disagree. If any of those places is something you would have acted on, you have an answer rather than a measurement, and that is fine as long as nobody builds a trend out of it.
Then take the last analysis a chatbot gave you and check its evidence by hand: does every quote appear in the transcript word for word, and does each one actually support the sentence it was attached to. The second check is the one that fails, and it fails quietly.
Bring the transcript you already pasted somewhere. Send us one conversation you have already run through a chatbot and we will mark which of its conclusions survive a word-for-word check against the recording, and which quietly do not. One transcript, and the parts that hold.
- Twelve of the top search results on this are vendor pages. Not one of them reports a measurement.
- A model asked for proof produces proof. The quote is often real and often does not support the claim it was attached to.
- At temperature zero, with a fixed seed, the same question can swing more than thirty points of accuracy between runs.
- Past a certain volume, supplying more context makes the answer worse than supplying none at all.
- The evidence is usually in the document. One pass is just bad at finding it.
A live demo of our products and a real conversation about the growth problem you need to solve.
Book a demo