Extended MBPP with additional test cases. Uses 399 hand-verified problems from MBPP-sanitized.
Pass@1 is the reported evaluation metric for MBPP+. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Qwen2.5-Coder-32B | verified | 76.4 | 2026 | Source ↗ | Looks wrong? |
| 02 | DeepSeek-V3 | verified | 73 | 2026 | Source ↗ | Looks wrong? |
| 03 | GPT-4o | verified | 71.2 | 2026 | Source ↗ | Looks wrong? |
| 04 | DeepSeek-Coder-33B | verified | 66 | 2026 | Source ↗ | Looks wrong? |