974 crowd-sourced Python programming problems suitable for beginners. Covers programming fundamentals and standard library.
Pass@1 is the reported evaluation metric for MBPP. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Claude 3.5 Sonnet (Oct 2024) | verified | 91 | 2026 | Source ↗ | Looks wrong? |
| 02 | Qwen2.5-Coder-32B-Instruct | verified | 90.2 | 2026 | Source ↗ | Looks wrong? |
| 03 | DeepSeek-Coder-V2-Instruct | verified | 89.4 | 2026 | Source ↗ | Looks wrong? |
| 04 | claude-35-sonnet | vendor | 89.2 | 2026 | Source ↗ | Looks wrong? |
| 05 | gpt-4o | vendor | 87.8 | 2026 | Source ↗ | Looks wrong? |
| 06 | GPT-4o (Aug 2024) | verified | 86.8 | 2026 | Source ↗ | Looks wrong? |
| 07 | Qwen2.5-Coder-7B-Instruct | verified | 83.5 | 2026 | Source ↗ | Looks wrong? |
| 08 | Codestral 22B v0.1 | verified | 78.2 | 2026 | Source ↗ | Looks wrong? |
| 09 | Llama 4 Maverick (17B-128E) | verified | 77.6 | 2026 | Source ↗ | Looks wrong? |
| 10 | DeepSeek-V3 | verified | 75.4 | 2026 | Source ↗ | Looks wrong? |
| 11 | Gemma 3 27B IT | verified | 74.4 | 2026 | Source ↗ | Looks wrong? |
| 12 | Gemma 3 12B IT | verified | 73 | 2026 | Source ↗ | Looks wrong? |
| 13 | Llama 4 Scout (17B-16E) | verified | 67.8 | 2026 | Source ↗ | Looks wrong? |
| 14 | Gemma 3 4B IT | verified | 63.2 | 2026 | Source ↗ | Looks wrong? |