# Whisper on 34 minutes of real recordings: word error rate by condition (baselines, October 2026)

> Four sizes of OpenAI's open-source Whisper model, run on a laptop CPU over 23 openly licensed English clips with human transcripts: clean and harder read speech, the same speech in restaurant noise, accented conversations and four-person meetings. Word error rate by condition, with intervals, every transcript and the code to reproduce it. These are baselines for comparing tools, not a ranking of products.

By TranscribeBench. Updated 8 October 2026. Canonical URL: https://transcribebench.com/whisper-baseline-benchmark/

**What this is.** A first measurement of how much the recording changes transcription accuracy, using open-source models anyone can run, on 23 clips (33.6 minutes) of openly licensed English speech with human transcripts. It is a yardstick for the commercial tools to be tested next, not a ranking of products: no commercial tool could be tested in this pilot within our rules ([why](#commercial-tools)).

**The short answer.** The recording matters more than the model. The largest model here, Whisper medium.en, got 1.9% of words wrong on clean audiobook speech and 22.6% in four-person meetings. Going from the smallest model to the largest cut the error rate on meetings from 27.3% to 22.6%; going from clean speech to meetings multiplied it 12 times.

## Results

<section class="benchmark-data" data-block="data" markdown="1">

### Word error rate by condition (lower is better)

Errors divided by reference words over all clips in the condition, after the [standard normalization](/wer-calculator/#normalization-option-by-option). The range under each overall figure is a 95% interval from resampling the clips.

<div class="table-wrap" markdown="1">

| Model | Clean read speech | Harder read speech | Restaurant noise, 5 dB | Accented conversation | Meetings | All 23 clips | Character error rate | Speed |
|---|---|---|---|---|---|---|---|---|
| Whisper tiny.en | 5.5% | 5.3% | 12.8% | 17.6% | 27.3% | **16.0%**<br><small>10.6% to 21.1%</small> | 10.0% | 17x real time |
| Whisper base.en | 4.0% | 5.3% | 7.7% | 16.4% | 25.4% | **14.1%**<br><small>8.9% to 18.9%</small> | 8.9% | 10x real time |
| Whisper small.en | 2.4% | 3.4% | 3.2% | 12.9% | 25.6% | **12.0%**<br><small>6.4% to 17.2%</small> | 8.2% | 3x real time |
| Whisper medium.en | 1.9% | 3.2% | 3.7% | 13.6% | 22.6% | **11.3%**<br><small>6.4% to 15.8%</small> | 7.5% | 2x real time |
| Clips | 5 | 4 | 5 | 5 | 4 | 23 | | |

</div>

Speed is audio length divided by processing time on a 2019 laptop CPU (Intel(R) Core(TM) i9-9880H CPU @ 2.30GHz, 8 of 16 threads), int8. A GPU is many times faster.

### What kind of errors

For Whisper medium.en, per condition: substitutions (a wrong word), deletions (a missed word) and insertions (an extra word).

<div class="table-wrap" markdown="1">

| Condition | Reference words | Substitutions | Deletions | Insertions |
|---|---|---|---|---|
| Clean read speech | 884 | 11 | 3 | 3 |
| Harder read speech | 685 | 17 | 2 | 3 |
| Restaurant noise, 5 dB | 884 | 21 | 9 | 3 |
| Accented conversation | 1,306 | 35 | 133 | 10 |
| Meetings | 1,554 | 68 | 261 | 22 |

</div>

</section>

## What the numbers say

- **Clean read speech is close to solved.** On the clean LibriSpeech clips the models got between 1.9% and 5.5% of words wrong. Treat these as a best case: LibriSpeech comes from public-domain LibriVox audiobooks, and Whisper's training data, which OpenAI has not published, may include some of the same recordings.
- **Noise costs accuracy, and small models pay most.** The noisy clips are the five clean clips with a restaurant recording mixed in at 5 dB below the speech. Whisper tiny.en went from 5.5% to 12.8%; Whisper medium.en from 1.9% to 3.7%.
- **Conversations are harder than reading,** and meetings hardest. In the meeting clips most errors are words left out: 261 of the 351 errors for Whisper medium.en are deletions, and 143 of those missed words come in stretches of five or more. Reading the alignments, the stretches are quick exchanges ("thank you", "okay, I'll have a go"), asides, and the ends of turns where someone else starts talking. A meeting transcript can read well and still leave out whole remarks.
- **Accented conversation sits in between.** The EdAcc clips are relaxed video-call conversations between friends; besides English, the speakers' first languages are Bulgarian, Lithuanian, Hindi, Konkani, Marathi, Romanian, Sinhalese, Mandarin, Shona, Catalan and Spanish. Five clips cannot tell you how any one accent fares; the [per-clip table](#every-clip) shows the spread.
- **Five clips per condition is a small sample.** The intervals are wide, and differences of a point or two between models on one condition are within the noise: Whisper small.en did better than medium.en on the noisy clips and the accented conversations. The gap between conditions is not.

## Every clip

<details markdown="1">
<summary>Word error rate for every clip and model</summary>

<div class="table-wrap" markdown="1">

| Clip | Condition | Length | Reference words | tiny.en | base.en | small.en | medium.en | 
|---|---|---|---|---|---|---|---|
| ls-clean-1 | clean-read | 68 s | 141 | 3.5% | 1.4% | 0.7% | 1.4% | 
| ls-clean-2 | clean-read | 73 s | 231 | 7.9% | 5.4% | 5.4% | 5.0% | 
| ls-clean-3 | clean-read | 72 s | 194 | 6.2% | 3.1% | 1.5% | 0.0% | 
| ls-clean-4 | clean-read | 61 s | 151 | 5.4% | 7.4% | 2.0% | 0.7% | 
| ls-clean-5 | clean-read | 68 s | 161 | 3.1% | 1.9% | 0.6% | 1.2% | 
| ls-other-1 | hard-read | 72 s | 190 | 8.9% | 7.8% | 5.2% | 4.7% | 
| ls-other-2 | hard-read | 67 s | 178 | 4.4% | 3.3% | 2.8% | 3.3% | 
| ls-other-3 | hard-read | 66 s | 167 | 4.1% | 3.5% | 2.3% | 2.9% | 
| ls-other-4 | hard-read | 69 s | 141 | 2.8% | 6.4% | 2.8% | 1.4% | 
| noisy-1 | noisy | 68 s | 141 | 11.3% | 5.0% | 0.7% | 4.3% | 
| noisy-2 | noisy | 73 s | 231 | 22.9% | 13.3% | 7.5% | 7.1% | 
| noisy-3 | noisy | 72 s | 194 | 9.3% | 6.7% | 2.6% | 2.6% | 
| noisy-4 | noisy | 61 s | 151 | 11.5% | 8.8% | 2.0% | 1.4% | 
| noisy-5 | noisy | 68 s | 161 | 4.3% | 1.9% | 0.6% | 1.9% | 
| ami-1 | meeting | 121 s | 290 | 30.3% | 30.3% | 32.2% | 22.1% | 
| ami-2 | meeting | 121 s | 220 | 23.7% | 22.3% | 28.8% | 27.9% | 
| ami-3 | meeting | 121 s | 287 | 13.0% | 13.0% | 10.2% | 13.0% | 
| ami-4 | meeting | 173 s | 700 | 32.7% | 29.2% | 27.9% | 25.0% | 
| edacc-1 | accented | 97 s | 196 | 19.6% | 14.1% | 15.1% | 14.6% | 
| edacc-2 | accented | 97 s | 328 | 16.8% | 13.1% | 11.9% | 12.8% | 
| edacc-3 | accented | 93 s | 182 | 13.4% | 19.8% | 7.0% | 7.6% | 
| edacc-4 | accented | 122 s | 323 | 12.3% | 14.2% | 9.9% | 9.6% | 
| edacc-5 | accented | 111 s | 284 | 26.2% | 22.5% | 19.6% | 22.5% | 

</div>

</details>

Every reference transcript and every model transcript is in [benchmark-transcripts-baselines.json](/tools/wer-calculator/benchmark-transcripts-baselines.json). Paste any pair into the [WER calculator](/wer-calculator/) with the Standard preset and you get the same number: it is the same code.

## The test set

23 clips, 33.6 minutes, chosen by fixed rules written down before any model was run ([methodology](/methodology/#the-test-set)):

| Condition | Clips | Source | What it is |
|---|---|---|---|
| Clean read speech | 5 | LibriSpeech test-clean | One reader per clip (3 women, 2 men), about 70 seconds of consecutive sentences from an audiobook chapter. |
| Harder read speech | 4 | LibriSpeech test-other | Readers the corpus authors rated harder to recognize (2 women, 2 men). |
| Restaurant noise, 5 dB | 5 | The clean clips plus a public-domain restaurant recording | Speech 5 dB louder than the noise, measured over the whole clip. |
| Accented conversation | 5 | EdAcc test set | 90 to 120 seconds of a two-person video call between friends. |
| Meetings | 4 | AMI Meeting Corpus test set | Two to three minutes of a 3-4 person meeting, all headsets mixed together. |

## Commercial tools

The plan was to run commercial tools through the same clips. None could be tested in this pilot within the rules: free tiers only, no accounts, nothing that the tool's terms forbid (benchmarking, competitive analysis, automated access), and no working around a bot check or a limit. Most transcription apps need an account even for a free trial; several tools' terms forbid benchmarking or automated access outright; one free no-account tool sits behind a bot check; and one returned only the first part of a transcript without an account. The tool-by-tool record, with the terms quoted, will be published with the first commercial results, after each tested vendor has seen its numbers. Until then the [cost calculator](/transcription-cost-calculator/) compares their prices, and the [WER calculator](/wer-calculator/) lets you test any tool on your own audio.

## How it was run

- **Models.** OpenAI Whisper tiny.en, base.en, small.en, medium.en (English-only checkpoints, MIT license), in the CTranslate2 conversions run by faster-whisper 1.2.1 with CTranslate2 4.8.2, int8 on the CPU, English, beam size 5, no voice-activity filter, otherwise the library defaults. Nothing was tuned on these clips. Run on 8 October 2026.
- **Scoring.** The same JavaScript as the [WER calculator](/wer-calculator/), with the Standard normalization, rates pooled over clips; every clip was re-scored with jiwer 4.0.0 as a check, with no difference.
- **Reproduce it.** `data/transcribe/fetch_samples.py` downloads the corpora and rebuilds every clip byte for byte (SHA-256 checked); `run_baselines.py` runs the models; `score.py` scores. The [sample list](/tools/wer-calculator/benchmark-samples.yaml) has every clip's exact source utterances or time window.

## Credits

The clips are cut from openly licensed research corpora; the audio is not redistributed here.

- **LibriSpeech** ASR corpus, © 2014 Vassil Panayotov, [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) (Panayotov, Chen, Povey and Khudanpur, "LibriSpeech: an ASR corpus based on public domain audio books", ICASSP 2015), from LibriVox recordings. [openslr.org/12](https://www.openslr.org/12). Changed: consecutive utterances joined; for the noisy clips, mixed with noise.
- **AMI Meeting Corpus**, [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/) (Carletta et al., 2005), headset mix and manual annotations v1.6.2. [groups.inf.ed.ac.uk/ami](https://groups.inf.ed.ac.uk/ami/corpus/). Changed: windows cut from meetings.
- **EdAcc**, The Edinburgh International Accents of English Corpus, Sanabria, Markl, Carmantini, Klejch, Bell and Bogoychev, University of Edinburgh, [doi:10.7488/ds/7914](https://doi.org/10.7488/ds/7914), [CC BY-SA 4.0](https://creativecommons.org/licenses/by-sa/4.0/). Changed: windows cut from conversations. The EdAcc reference transcripts republished in the transcripts file are shared under CC BY-SA 4.0.
- **Noise:** "Restaurant ambience" by stephan, public domain, from pdsounds.org via [Wikimedia Commons](https://commons.wikimedia.org/wiki/File:Restaurant_ambience.ogg).
