Skip to content

Accuracy benchmark

We measured our accuracy. Here is exactly how.

On the Google FLEURS benchmark, BriefVox makes about 5.3% word errors per 100 words across eight languages — down from 7.6% with our previous engine. Below: every number, how we got it, and where it stops applying.

average word error rate, 8 languages
5.3%
languages measured; 99 recognised
8
faster than real time (Parakeet step)
38×
previous engine, same test
7.6%

Word error rate by language

Lower is better. Each language is a 5–8 minute file of 30 joined FLEURS sentences, transcribed with automatic language detection — the way your uploads are processed.

Word error rate by language
LanguageEngine in productionWER todayPrevious engineTodayPrevious engine
EnglishWhisper large-v3-turbo5.4%3.8%
PolishParakeet-TDT-0.6B-v38.7%7.5%
GermanParakeet-TDT-0.6B-v37.4%8.8%
SpanishParakeet-TDT-0.6B-v33.0%4.5%
FrenchParakeet-TDT-0.6B-v34.9%8.1%
ItalianParakeet-TDT-0.6B-v33.5%13.7%
PortugueseParakeet-TDT-0.6B-v34.4%8.3%
RussianWhisper large-v3-turbo5.1%6.0%
Average5.3%7.6%

Measured on October 1, 2026 on the GPU that runs production transcription. Where the previous engine looks better (English, Polish), the gap did not survive a larger sample — see the re-run with 80 sentences below.

Why two engines

No single model won every language, so BriefVox routes each recording to the engine that measured best for its language. Whisper also identifies the language and covers everything outside Parakeet's 25 European languages.

EngineAverage WERSpeed
NVIDIA Parakeet-TDT-0.6B-v3ONNX, fp325.5%37.8× real time
OpenAI Whisper large-v3-turboint8, beam 56.3%10.8× real time
OpenAI Whisper large-v3int8, beam 10 — previous production setting7.6%4.4× real time

Close calls, re-checked

Thirty sentences is about 700 words per language, so a one-point gap is a handful of words. Where the leaders were close (English, German, Polish, Russian) we re-ran with 80 sentences. German and Polish came out as ties, so the 3× faster engine takes them; English and Russian stay on Whisper.

EngineENDEPLRU
Whisper large-v34.7%——6.1%
Whisper large-v3-turbo4.8%5.5%7.2%4.8%
Parakeet-TDT-0.6B-v35.3%5.4%7.1%6.5%

Method

Data

Google FLEURS test split for the eight interface languages. Sentences are joined with 0.6 s pauses into one long file, because real uploads are long — segmentation and voice detection are part of what is measured.

Level

Every sentence is brought to the same loudness (−23 dBFS) before joining, as our pipeline normalises loudness before recognition anyway.

Metric

Word error rate after the same normalisation for every engine: lowercase, punctuation and symbols removed. Numbers are not reconciled (“25” vs “twenty-five”), which costs every engine a little.

Speed

Seconds of audio per second of processing on the production GPU (NVIDIA GTX 1060 6 GB), model loading excluded.

What these numbers don't tell you

  • FLEURS is clearly read speech. Noisy rooms, cross-talk, phone audio and strong accents produce more errors — you can fix any word in the editor.
  • About 700 words per language: differences under a point are within noise.
  • Word error rate counts words, not speaker labels or timestamps; those are produced by separate alignment and diarization steps.
  • Recognition is measured on its own. A full job also uploads, aligns words, finds speakers and saves the result, so end-to-end time is longer than the speed figure.

What happens to your file

  1. 1Whisper large-v3-turbo detects the language from several 30-second windows across the recording, not just the intro.
  2. 2The recording goes to the engine that measured best for that language; if Parakeet fails or doesn't know the language, Whisper takes over.
  3. 3Every word is timed with forced alignment, and speakers are separated with pyannote.

Reproduce it

The benchmark scripts download FLEURS, run each engine and score the output. These are the exact commands behind the table:

python fetch_fleurs.py --out data --per-lang 30
python run_bench.py --engine onnx --model nemo-parakeet-tdt-0.6b-v3 \
    --repo istupakov/parakeet-tdt-0.6b-v3-onnx --kinds long --out results/parakeet.jsonl
python run_bench.py --engine whisperx --model large-v3-turbo --beam 5 --kinds long \
    --out results/turbo.jsonl
python score.py results/*.jsonl

Try it on your own recording

20 free minutes every month, no card. Your audio is the benchmark that matters.

Start free