Last updated: September 2026
Every transcription vendor advertises an accuracy figure, most of them somewhere in the high nineties, and almost nobody who quotes one back at you can say what it was measured on. This article is about that gap: what the number means, why your own recordings will not reproduce it, which errors actually cost you something, and how to run a test that takes about twenty minutes and tells you more than any vendor page will.
If what you want is the roundup of tools by category — voice, music, podcast production, transcription — that is a separate page: the best AI audio tools. This one assumes you have a shortlist and need to find out whether any of them survives contact with your actual audio.
What an Accuracy Percentage Is Actually Measuring
The standard metric in speech recognition is word error rate, usually written WER. It is defined as the number of substitutions, insertions and deletions needed to turn the machine's transcript into the reference transcript, divided by the number of words in the reference. An accuracy claim of "95%" is normally a WER of 5% restated in the more flattering direction.
Three things follow from that definition, and all three are routinely lost when the number is quoted:
- It is entirely relative to a reference transcript that somebody produced by hand, on a particular recording. Change the recording and the number changes. There is no absolute accuracy.
- Every error weighs the same. Transcribing "the" as "a" and transcribing a client's name as a different word count identically. So do "did" and "did not".
- It says nothing about the parts you will actually complain about — who was speaking, where the sentence ended, whether the timestamps line up.
None of this makes the metric useless. It makes it a laboratory measurement quoted as a field guarantee, which is a different problem, and the fix is to measure it yourself on audio that looks like yours.
Why the Number Falls Apart on Your Recordings
Vendor benchmarks tend to use clean audio: a single speaker, close microphone, quiet room, general-interest vocabulary. Real recordings differ from that in five specific ways, each of which degrades output independently.
- Overlapping speech. Two people talking at once is the single hardest case, and it is the normal condition in a lively meeting. Systems handle it by dropping one speaker, merging both into nonsense, or guessing.
- Accent and second-language speech. Recognition quality varies substantially by accent, and by how far the speaker's pronunciation sits from whatever dominated the training data. This is the same structural bias that shows up in AI text detection, where the tools misfire hardest on competent non-native writing — we went through the published evidence on that in our piece on AI detector accuracy.
- Domain vocabulary. Drug names, legal terms, ticker symbols, internal project code names, the surname of everyone on the call. These are the words that carry the meaning and they are the words most likely to be wrong, because they are rare in general training data.
- The microphone. A laptop's built-in microphone across a meeting room is a worse input than any software can fully rescue. The single cheapest accuracy improvement available to most people is a better microphone, not a better tool.
- Room and line noise. Air conditioning, keyboard clatter, a bad connection dropping syllables. Noise suppression helps and sometimes removes speech along with the noise.
The Errors That Actually Cost You Something
Because WER treats all errors alike, it is worth separating them yourself. In practice they fall into three tiers.
Harmless. Filler words dropped, "gonna" rendered as "going to", a slightly wrong article. Nobody is worse off. These make up the bulk of most error counts, which is why a transcript with a mediocre WER often reads perfectly well.
Annoying. Sentence boundaries in the wrong place, paragraphs that run on, speaker labels that swap mid-conversation. You can fix these, but you have to read the whole thing to find them, which erodes the reason you used a machine.
Expensive. Three kinds, and they are worth naming individually because they are the ones that survive review:
- Dropped negations. "We are not going to proceed" losing its "not" reverses the meaning of a sentence and reads as perfectly fluent English. Nothing in the transcript flags it.
- Numbers. Digits, dates, quantities and prices are frequently mis-heard and almost never look wrong. "Fifteen" and "fifty" are one phoneme apart.
- Proper nouns. A misspelt name in a transcript that gets circulated is a small professional embarrassment; in a record of a commitment, it is worse.
When you evaluate a tool, weight these three heavily and largely ignore the filler-word errors. A tool with a worse headline rate that never drops a negation is the better tool.
The Twenty-Minute Test
This is the part worth doing, and it is not difficult.
- Pick a genuinely representative recording. Not your best one. A normal meeting with your normal people, your normal microphone and your normal jargon. Five to ten minutes is enough.
- Transcribe one minute of it by hand — properly, every word. This is the tedious part and it is what makes the test real. You now have a reference nobody can argue with.
- Run the same file through each shortlisted tool on the same day. Same file, not a re-recording.
- Compare against your minute and count errors in the three expensive categories separately from everything else.
- Then read the rest of the machine transcript looking specifically at names, numbers and negations, checking against your memory of the meeting.
- Check the non-text output. Are speaker labels consistent from start to finish? Do timestamps land on the right words when you click them? Can you export in a format you can actually use later?
Comparative testing is the point. "Is this accurate" is a question nobody can answer honestly. "Is this one better than that one on my audio" is answerable in an afternoon.



