wmt23 is a state-of-the-art machine learning benchmark indexed on Codesota. This page tracks published model results, top scores per metric, and the SOTA timeline for wmt23.
Comet is the reported evaluation metric for wmt23. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | GPT-4 | verified | 84.1 | 2026 | Source ↗ | Looks wrong? |
| 02 | Google Translate | verified | 83.8 | 2026 | Source ↗ | Looks wrong? |
| 03 | DeepL | verified | 83.5 | 2026 | Source ↗ | Looks wrong? |
| 04 | NLLB-3.3B | verified | 81.6 | 2026 | Source ↗ | Looks wrong? |