AI ToolsJune 3, 2026

AI Transcription Accuracy: What the Percentage on the Box Actually Measures

An audio waveform built from rows of thin vertical bars of varying height, lit blue and magenta against black

Reviewed by NorwegianSpark Editorial | NorwegianSpark SA

Written with AI assistance and reviewed by the NorwegianSpark SA editorial team.

Last updated: September 2026

Every transcription vendor advertises an accuracy figure, most of them somewhere in the high nineties, and almost nobody who quotes one back at you can say what it was measured on. This article is about that gap: what the number means, why your own recordings will not reproduce it, which errors actually cost you something, and how to run a test that takes about twenty minutes and tells you more than any vendor page will.

If what you want is the roundup of tools by category — voice, music, podcast production, transcription — that is a separate page: the best AI audio tools. This one assumes you have a shortlist and need to find out whether any of them survives contact with your actual audio.

What an Accuracy Percentage Is Actually Measuring

The standard metric in speech recognition is word error rate, usually written WER. It is defined as the number of substitutions, insertions and deletions needed to turn the machine's transcript into the reference transcript, divided by the number of words in the reference. An accuracy claim of "95%" is normally a WER of 5% restated in the more flattering direction.

Three things follow from that definition, and all three are routinely lost when the number is quoted:

  • It is entirely relative to a reference transcript that somebody produced by hand, on a particular recording. Change the recording and the number changes. There is no absolute accuracy.
  • Every error weighs the same. Transcribing "the" as "a" and transcribing a client's name as a different word count identically. So do "did" and "did not".
  • It says nothing about the parts you will actually complain about — who was speaking, where the sentence ended, whether the timestamps line up.

None of this makes the metric useless. It makes it a laboratory measurement quoted as a field guarantee, which is a different problem, and the fix is to measure it yourself on audio that looks like yours.

Why the Number Falls Apart on Your Recordings

Vendor benchmarks tend to use clean audio: a single speaker, close microphone, quiet room, general-interest vocabulary. Real recordings differ from that in five specific ways, each of which degrades output independently.

  • Overlapping speech. Two people talking at once is the single hardest case, and it is the normal condition in a lively meeting. Systems handle it by dropping one speaker, merging both into nonsense, or guessing.
  • Accent and second-language speech. Recognition quality varies substantially by accent, and by how far the speaker's pronunciation sits from whatever dominated the training data. This is the same structural bias that shows up in AI text detection, where the tools misfire hardest on competent non-native writing — we went through the published evidence on that in our piece on AI detector accuracy.
  • Domain vocabulary. Drug names, legal terms, ticker symbols, internal project code names, the surname of everyone on the call. These are the words that carry the meaning and they are the words most likely to be wrong, because they are rare in general training data.
  • The microphone. A laptop's built-in microphone across a meeting room is a worse input than any software can fully rescue. The single cheapest accuracy improvement available to most people is a better microphone, not a better tool.
  • Room and line noise. Air conditioning, keyboard clatter, a bad connection dropping syllables. Noise suppression helps and sometimes removes speech along with the noise.

The Errors That Actually Cost You Something

Because WER treats all errors alike, it is worth separating them yourself. In practice they fall into three tiers.

Harmless. Filler words dropped, "gonna" rendered as "going to", a slightly wrong article. Nobody is worse off. These make up the bulk of most error counts, which is why a transcript with a mediocre WER often reads perfectly well.

Annoying. Sentence boundaries in the wrong place, paragraphs that run on, speaker labels that swap mid-conversation. You can fix these, but you have to read the whole thing to find them, which erodes the reason you used a machine.

Expensive. Three kinds, and they are worth naming individually because they are the ones that survive review:

  • Dropped negations. "We are not going to proceed" losing its "not" reverses the meaning of a sentence and reads as perfectly fluent English. Nothing in the transcript flags it.
  • Numbers. Digits, dates, quantities and prices are frequently mis-heard and almost never look wrong. "Fifteen" and "fifty" are one phoneme apart.
  • Proper nouns. A misspelt name in a transcript that gets circulated is a small professional embarrassment; in a record of a commitment, it is worse.

When you evaluate a tool, weight these three heavily and largely ignore the filler-word errors. A tool with a worse headline rate that never drops a negation is the better tool.

The Twenty-Minute Test

This is the part worth doing, and it is not difficult.

  1. Pick a genuinely representative recording. Not your best one. A normal meeting with your normal people, your normal microphone and your normal jargon. Five to ten minutes is enough.
  2. Transcribe one minute of it by hand — properly, every word. This is the tedious part and it is what makes the test real. You now have a reference nobody can argue with.
  3. Run the same file through each shortlisted tool on the same day. Same file, not a re-recording.
  4. Compare against your minute and count errors in the three expensive categories separately from everything else.
  5. Then read the rest of the machine transcript looking specifically at names, numbers and negations, checking against your memory of the meeting.
  6. Check the non-text output. Are speaker labels consistent from start to finish? Do timestamps land on the right words when you click them? Can you export in a format you can actually use later?

Comparative testing is the point. "Is this accurate" is a question nobody can answer honestly. "Is this one better than that one on my audio" is answerable in an afternoon.

The Summary Problem, Which Is Worse Than the Transcript Problem

Most meeting tools now produce a summary and a list of action items rather than only a transcript, and this is where the risk profile changes sharply.

A transcript error is visible: the sentence reads oddly and you go and check. A summary error is invisible, because the summary is the only artefact anyone reads. If the model missed a caveat, compressed a disagreement into an agreement, or attributed a commitment to the wrong person, there is nothing in the output that looks wrong. It reads as a clean, confident set of decisions.

The practical rule: for anything consequential, keep the transcript and treat the summary as an index into it rather than a replacement for it. If a summary says a decision was made, the decision should be traceable to a timestamp. Tools that let you jump from a summary line to the moment in the audio are considerably more trustworthy in use than tools that only hand you the paragraph, regardless of which produces better prose.

Recording People: The Part That Is Not a Software Question

Recording a conversation carries legal obligations, and they differ by jurisdiction, by whether it is a phone call or a meeting, by whether all parties or only one need to consent, and by what you then do with the recording. We are not going to print a rule here, because a wrong one is genuinely dangerous and the correct one depends on where everyone on the call is sitting.

What is portable is the set of questions to resolve before you switch anything on: who needs to consent and how is that consent recorded; where is the audio stored and for how long; who can access the transcript afterwards; and what happens to all of it if someone asks for their data to be deleted. Your own legal or compliance advice is the source for the answers. Most transcription vendors document retention and deletion behaviour, and that documentation is the thing to read rather than the marketing page.

The social layer matters too, separately from the legal one. Announcing that a meeting is being transcribed changes how people speak in it, and that is a real cost in some conversations and a real benefit in others.

What to Check Before You Commit

  • What is the input unit — minutes, hours, files, seats — and what happens when you exceed it?
  • Can you export the transcript in a plain format you could still read if the vendor disappeared?
  • Does it support the languages your team actually uses, including mixed-language conversations?
  • Can you supply a custom vocabulary of names and terms? This is the single most effective accuracy feature for domain-heavy audio.
  • Is your audio used to train their models, and can that be switched off? Look for the setting, not the reassurance.
  • Does it join meetings as a visible participant, and is that acceptable to your clients?

The Honest Counter-Argument

There are still cases where a human transcriber is the correct answer, and the AI tools do not close them. Anything intended as a legal record, anything where a certified transcript is required, and anything where a single dropped negation would be materially damaging all belong to a human with an accountability chain. The value of a professional transcriber is not only accuracy but the fact that someone is answerable for it.

The opposite point also holds, and it is the one people forget when comparing against perfection: the realistic alternative to a mediocre automatic transcript is usually no transcript at all, plus one person's handwritten notes. Measured against that, even an imperfect machine transcript is a large improvement, and the search function alone often justifies it.

Tools We Hold, and What We Are Not Claiming

We have affiliate relationships with several transcription products, and we are not going to attach an accuracy figure to any of them, because we have not run a controlled benchmark and quoting the vendors' own numbers would be repeating an unverified claim. What we can honestly say is what each is built for: Transkriptor is an audio and video transcription tool aimed at turning recordings into text you can edit, and Notta is a meeting notetaker that transcribes and summarises calls and interviews into searchable text. For teams where the recording itself is the problem rather than the transcription, EyesOn handles the video meeting layer.

Run the twenty-minute test on your own audio before paying for any of them, including ours. It is the only comparison that describes your situation.

Disclosure: this article contains affiliate links. If you sign up through them we may earn a commission at no extra cost to you. It does not change what we recommend and no vendor has paid for a position here.

Related Articles

Continue reading

Continue in this collection