Codesota · Benchmark · HumanEval+Home/Leaderboards/Code & Software Engineering/Code Generation/HumanEval+
Unknown

HumanEval+.

Extended HumanEval with 80x more test cases. Tests code robustness and edge case handling.

Paper ↗Leaderboard ↓Lineage
§ 01 · Leaderboard

Results by metric.

Found a wrong score or missing run?
Use row edits to send a sourced correction into moderation.
Add / edit result ↗Report issue ↗

pass@1

Pass@1 is the reported evaluation metric for HumanEval+. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for pass@1verifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Qwen2.5-Coder-32B
Fetched from CodeSOTA API on 2026-04-20
verified87.22026Source ↗Looks wrong?
02DeepSeek-V3
Fetched from CodeSOTA API on 2026-04-20
verified86.62026Source ↗Looks wrong?
03GPT-4o
Fetched from CodeSOTA API on 2026-04-20
verified862026Source ↗Looks wrong?
04DeepSeek-Coder-V2
Fetched from CodeSOTA API on 2026-04-20
verified82.32026Source ↗Looks wrong?
05DeepSeek-Coder-33B
Fetched from CodeSOTA API on 2026-04-20
verified752026Source ↗Looks wrong?
Lineage

HumanEval+ in context.

See full coding benchmarks lineage →
This benchmark (1)
active2023-05
HumanEval+
§ 04 · Submit a result

Add to the leaderboard.

Submit a Result

Sign in to submit benchmark results for HumanEval+.

Sign in
← Back to Code Generation