What this is. A first measurement of how much the recording changes transcription accuracy, using open-source models anyone can run, on 23 clips (33.6 minutes) of openly licensed English speech with human transcripts. It is a yardstick for the commercial tools to be tested next, not a ranking of products: no commercial tool could be tested in this pilot within our rules (why).

The short answer. The recording matters more than the model. The largest model here, Whisper medium.en, got 1.9% of words wrong on clean audiobook speech and 22.6% in four-person meetings. Going from the smallest model to the largest cut the error rate on meetings from 27.3% to 22.6%; going from clean speech to meetings multiplied it 12 times.

Results

Word error rate by condition (lower is better)

Errors divided by reference words over all clips in the condition, after the standard normalization. The range under each overall figure is a 95% interval from resampling the clips.

Model Clean read speech Harder read speech Restaurant noise, 5 dB Accented conversation Meetings All 23 clips Character error rate Speed
Whisper tiny.en 5.5% 5.3% 12.8% 17.6% 27.3% 16.0%
10.6% to 21.1%
10.0% 17x real time
Whisper base.en 4.0% 5.3% 7.7% 16.4% 25.4% 14.1%
8.9% to 18.9%
8.9% 10x real time
Whisper small.en 2.4% 3.4% 3.2% 12.9% 25.6% 12.0%
6.4% to 17.2%
8.2% 3x real time
Whisper medium.en 1.9% 3.2% 3.7% 13.6% 22.6% 11.3%
6.4% to 15.8%
7.5% 2x real time
Clips 5 4 5 5 4 23

Speed is audio length divided by processing time on a 2019 laptop CPU (Intel(R) Core(TM) i9-9880H CPU @ 2.30GHz, 8 of 16 threads), int8. A GPU is many times faster.

What kind of errors

For Whisper medium.en, per condition: substitutions (a wrong word), deletions (a missed word) and insertions (an extra word).

Condition Reference words Substitutions Deletions Insertions
Clean read speech 884 11 3 3
Harder read speech 685 17 2 3
Restaurant noise, 5 dB 884 21 9 3
Accented conversation 1,306 35 133 10
Meetings 1,554 68 261 22

What the numbers say

Every clip

Word error rate for every clip and model
Clip Condition Length Reference words tiny.en base.en small.en medium.en
ls-clean-1 clean-read 68 s 141 3.5% 1.4% 0.7% 1.4%
ls-clean-2 clean-read 73 s 231 7.9% 5.4% 5.4% 5.0%
ls-clean-3 clean-read 72 s 194 6.2% 3.1% 1.5% 0.0%
ls-clean-4 clean-read 61 s 151 5.4% 7.4% 2.0% 0.7%
ls-clean-5 clean-read 68 s 161 3.1% 1.9% 0.6% 1.2%
ls-other-1 hard-read 72 s 190 8.9% 7.8% 5.2% 4.7%
ls-other-2 hard-read 67 s 178 4.4% 3.3% 2.8% 3.3%
ls-other-3 hard-read 66 s 167 4.1% 3.5% 2.3% 2.9%
ls-other-4 hard-read 69 s 141 2.8% 6.4% 2.8% 1.4%
noisy-1 noisy 68 s 141 11.3% 5.0% 0.7% 4.3%
noisy-2 noisy 73 s 231 22.9% 13.3% 7.5% 7.1%
noisy-3 noisy 72 s 194 9.3% 6.7% 2.6% 2.6%
noisy-4 noisy 61 s 151 11.5% 8.8% 2.0% 1.4%
noisy-5 noisy 68 s 161 4.3% 1.9% 0.6% 1.9%
ami-1 meeting 121 s 290 30.3% 30.3% 32.2% 22.1%
ami-2 meeting 121 s 220 23.7% 22.3% 28.8% 27.9%
ami-3 meeting 121 s 287 13.0% 13.0% 10.2% 13.0%
ami-4 meeting 173 s 700 32.7% 29.2% 27.9% 25.0%
edacc-1 accented 97 s 196 19.6% 14.1% 15.1% 14.6%
edacc-2 accented 97 s 328 16.8% 13.1% 11.9% 12.8%
edacc-3 accented 93 s 182 13.4% 19.8% 7.0% 7.6%
edacc-4 accented 122 s 323 12.3% 14.2% 9.9% 9.6%
edacc-5 accented 111 s 284 26.2% 22.5% 19.6% 22.5%

Every reference transcript and every model transcript is in benchmark-transcripts-baselines.json. Paste any pair into the WER calculator with the Standard preset and you get the same number: it is the same code.

The test set

23 clips, 33.6 minutes, chosen by fixed rules written down before any model was run (methodology):

Condition Clips Source What it is
Clean read speech 5 LibriSpeech test-clean One reader per clip (3 women, 2 men), about 70 seconds of consecutive sentences from an audiobook chapter.
Harder read speech 4 LibriSpeech test-other Readers the corpus authors rated harder to recognize (2 women, 2 men).
Restaurant noise, 5 dB 5 The clean clips plus a public-domain restaurant recording Speech 5 dB louder than the noise, measured over the whole clip.
Accented conversation 5 EdAcc test set 90 to 120 seconds of a two-person video call between friends.
Meetings 4 AMI Meeting Corpus test set Two to three minutes of a 3-4 person meeting, all headsets mixed together.

Commercial tools

The plan was to run commercial tools through the same clips. None could be tested in this pilot within the rules: free tiers only, no accounts, nothing that the tool's terms forbid (benchmarking, competitive analysis, automated access), and no working around a bot check or a limit. Most transcription apps need an account even for a free trial; several tools' terms forbid benchmarking or automated access outright; one free no-account tool sits behind a bot check; and one returned only the first part of a transcript without an account. The tool-by-tool record, with the terms quoted, will be published with the first commercial results, after each tested vendor has seen its numbers. Until then the cost calculator compares their prices, and the WER calculator lets you test any tool on your own audio.

How it was run

Credits

The clips are cut from openly licensed research corpora; the audio is not redistributed here.

Also available as Markdown.