VoiceBench is a multi-facet evaluation suite for LLM-based voice assistants, covering general knowledge, instruction following, safety refusal, and robustness to speaker accents and background noise across diverse speech inputs.
Overall Score is the reported evaluation metric for VoiceBench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Ultravox-GLM-4P7 | verified | 88.86 | 2026 | Source ↗ | Looks wrong? |
| 02 | Whisper-v3-large + GPT-4o (cascade) | verified | 87.8 | 2026 | Source ↗ | Looks wrong? |
| 03 | GPT-4o-Audio | verified | 86.75 | 2026 | Source ↗ | Looks wrong? |
| 04 | Whisper-v3-large + LLaMA-3.1-8B (cascade) | verified | 77.48 | 2026 | Source ↗ | Looks wrong? |
| 05 | Kimi-Audio | verified | 76.91 | 2026 | Source ↗ | Looks wrong? |
| 06 | MiniCPM-o | verified | 71.23 | 2026 | Source ↗ | Looks wrong? |
| 07 | VITA-1.5 | verified | 64.53 | 2026 | Source ↗ | Looks wrong? |
| 08 | Qwen2-Audio | verified | 55.8 | 2026 | Source ↗ | Looks wrong? |
| 09 | LLaMA-Omni | verified | 41.12 | 2026 | Source ↗ | Looks wrong? |
| 10 | VITA-1.0 | verified | 36.43 | 2026 | Source ↗ | Looks wrong? |
| 11 | Mini-Omni2 | verified | 33.49 | 2026 | Source ↗ | Looks wrong? |
| 12 | Mini-Omni | verified | 30.42 | 2026 | Source ↗ | Looks wrong? |
| 13 | Moshi | verified | 29.51 | 2026 | Source ↗ | Looks wrong? |