Codesota · Speech · Speech Recognition · VoxPopuliTasks/Speech/Speech Recognition
Speech Recognition · benchmark dataset · 2021 · MULTILINGUAL

VoxPopuli Multilingual Speech Corpus.

VoxPopuli is a large-scale multilingual speech corpus derived from European Parliament event recordings, providing labelled ASR data for 18 European languages plus large quantities of unlabelled audio for self-supervised pre-training.

Paper ↗Submit a result ↵
§ 01 · Leaderboard

Best published scores.

55 results indexed across 1 metric. Shaded row marks current SOTA; ties broken by submission date.


Primary
wer · lower is better
wer· primary
55 rows
#ModelOrgSubmittedPaper / codewer
01Audio Flamingo 3—Jul 2025Audio Flamingo 3: Advancing Audio Intelligence with Full… · code5.55
02Canary-Qwen-2.5BOpenNVIDIAMar 2025Training and Inference Efficiency of Encoder-Decoder Spe…5.66
03Llama 3 Speech (70B)—Jul 2024The Llama 3 Herd of Models · code5.70
04Granite Speech 4.1 2BOpenIBMMay 2025Granite-speech: open-source speech-aware LLMs with stron…5.70
05Granite Speech 3.3 8BOpenIBMMay 2025Granite-speech: open-source speech-aware LLMs with stron…5.72
06Granite 4.0 1B SpeechOpenIBMMay 2025Granite-speech: open-source speech-aware LLMs with stron…5.84
07Granite Speech 3.3 2BOpenIBMMay 2025Granite-speech: open-source speech-aware LLMs with stron…5.93
08Phi-4 Multimodal InstructOpenMicrosoftMar 2025Phi-4-Mini Technical Report: Compact yet Powerful Multim…6.04
09Parakeet-rnnt-0.6b—May 2023Fast Conformer with Linearly Scalable Attention for Effi…6.08
10SYMPHONY-ASR—Jan 2026pwc-dump6.30
11Stt_en_fastconformer_ctc_large—May 2023Fast Conformer with Linearly Scalable Attention for Effi…6.34
12Qwen3-ASR-1.7BOpenAlibabaJan 2026Qwen3-ASR Technical Report · code6.35
13Stt_en_fastconformer_transducer_large—May 2023Fast Conformer with Linearly Scalable Attention for Effi…6.45
14Stt_en_conformer_ctc_large—May 2020Conformer: Convolution-augmented Transformer for Speech … · code6.83
15Asr-conformer-loquacious—Feb 2025pwc-dump6.89
16Parakeet-tdt_ctc-110m—Apr 2023Efficient Sequence Transduction by Jointly Predicting To… · code6.90
17Voxtral-Small-24B-2507OpenMistral AIJul 2025Voxtral6.96
18Qwen3-ASR-0.6BOpenAlibabaJan 2026Qwen3-ASR Technical Report · code7.07
19Parakeet-ctc-0.6b—May 2023Fast Conformer with Linearly Scalable Attention for Effi…7.07
20Whisper Large v2OpenOpenAIDec 2022Robust Speech Recognition via Large-Scale Weak Supervisi… · code7.48
21Whisper Large—Dec 2022Robust Speech Recognition via Large-Scale Weak Supervisi… · code7.76
22Lite-whisper-large-v3-fast—Feb 2025LiteASR: Efficient Automatic Speech Recognition with Low… · code7.79
23VibeVoice-ASR-HF—Jan 2026VIBEVOICE-ASR Technical Report8.01
24Whisper-medium.en—Dec 2022Robust Speech Recognition via Large-Scale Weak Supervisi… · code8.06
25Lite-whisper-large-v3-acc—Feb 2025LiteASR: Efficient Automatic Speech Recognition with Low… · code8.11
26Lite-whisper-large-v3-turbo-acc—Feb 2025LiteASR: Efficient Automatic Speech Recognition with Low… · code8.17
27Distil-large-v2—Nov 2023Distil-Whisper: Robust Knowledge Distillation via Large-… · code8.24
28Distil-large-v3—Nov 2023Distil-Whisper: Robust Knowledge Distillation via Large-… · code8.25
29Voxtral-Mini-4B-Realtime-2602OpenMistral AIFeb 2026Voxtral Realtime8.34
30Whisper-small.en—Dec 2022Robust Speech Recognition via Large-Scale Weak Supervisi… · code8.50
31Niagara-38m-batch.en—Feb 2026pwc-dump8.73
32Distil-small.en—Nov 2023Distil-Whisper: Robust Knowledge Distillation via Large-… · code8.79
33Distil-medium.en—Nov 2023Distil-Whisper: Robust Knowledge Distillation via Large-… · code9.00
34Stt_en_conformer_ctc_small—May 2020Conformer: Convolution-augmented Transformer for Speech … · code9.07
35Whisper Large v3OpenOpenAIDec 2022Robust Speech Recognition via Large-Scale Weak Supervisi… · code9.54
36Whisper-base.en—Dec 2022Robust Speech Recognition via Large-Scale Weak Supervisi… · code9.76
37Niagara-19m-batch.en—Feb 2026pwc-dump9.92
38Moonshine-base—Oct 2024Moonshine: Speech Recognition for Live Transcription and… · code10.84
39Whisper Large v3 TurboOpenOpenAIDec 2022Robust Speech Recognition via Large-Scale Weak Supervisi… · code11.87
40Whisper-tiny.en—Dec 2022Robust Speech Recognition via Large-Scale Weak Supervisi… · code12
41Asr-wav2vec2-librispeech—Jun 2021SpeechBrain: A General-Purpose Speech Toolkit · code13.72
42Moonshine-streaming-tiny—Jan 2026pwc-dump14.02
43Moonshine-tiny—Oct 2024Moonshine: Speech Recognition for Live Transcription and… · code14.11
44Mms-1b-all—May 2023Scaling Speech Technology to 1,000+ Languages · code17.63
45Wav2vec2-large-960h-lv60-self—Jun 2020wav2vec 2.0: A Framework for Self-Supervised Learning of… · code21.42
46Wav2vec2-conformer-rel-pos-large-960h-ft—Oct 2020fairseq S2T: Fast Speech-to-Text Modeling with fairseq · code22.39
47Hubert-xlarge-ls960-ft—Jun 2021HuBERT: Self-Supervised Speech Representation Learning b… · code22.47
48Wav2vec2-conformer-rope-large-960h-ft—Oct 2020fairseq S2T: Fast Speech-to-Text Modeling with fairseq · code22.61
49Hubert-large-ls960-ft—Jun 2021HuBERT: Self-Supervised Speech Representation Learning b… · code22.70
50Wav2vec2-large-robust-ft-libri-960h—Apr 2021Robust wav2vec 2.0: Analyzing Domain Shift in Self-Super… · code23.27
51Data2vec-audio-large-960h—Feb 2022data2vec: A General Framework for Self-supervised Learni… · code23.86
52Data2vec-audio-base-960h—Feb 2022data2vec: A General Framework for Self-supervised Learni… · code27.25
53Mms-1b-fl102—May 2023Scaling Speech Technology to 1,000+ Languages · code27.97
54wav2vec 2.0 Large (960h)OpenMeta AIJun 2020wav2vec 2.0: A Framework for Self-Supervised Learning of… · code30.09
55Wav2vec2-base-960h—Jun 2020wav2vec 2.0: A Framework for Self-Supervised Learning of… · code32.48
Fig 2 · Rows sorted by score within each metric. Shaded row marks SOTA. Dates reflect model or paper release where available, otherwise the date Codesota accessed the source.
§ 03 · Progress

5 steps
of state of the art.

Each row below marks a model that broke the previous record on wer. Intermediate submissions are kept in the leaderboard above; only SOTA-setting entries are re-listed here.

Lower scores win. Each subsequent entry improved upon the previous best.

SOTA line · wer
  1. May 16, 2020Stt_en_conformer_ctc_large6.83
  2. May 8, 2023Parakeet-rnnt-0.6b6.08
  3. Jul 31, 2024Llama 3 Speech (70B)5.70
  4. Mar 7, 2025Canary-Qwen-2.5BNVIDIA5.66
  5. Jul 10, 2025Audio Flamingo 35.55
Fig 3 · SOTA-setting models only. 5 entries span May 2020 → Jul 2025.
§ 04 · Literature

23 papers
tied to this benchmark.

Every paper below corresponds to at least one row in the leaderboard above. Click through for the arXiv preprint and, when available, the reference implementation.

§ 06 · Contribute

Have a score that beats
this table?

Submit a checkpoint and a reproduction script. We will run it, publish the score, and — if it takes the top — annotate the step on the progress chart with your name.

Submit a result ↵Read submission guide
What a submission needs
  • 01A public checkpoint or API endpoint
  • 02A reproduction script with frozen commit + seed
  • 03Declared evaluation environment (Python, deps)
  • 04One row per metric declared by this dataset
  • 05A contact so we can follow up on discrepancies