Codesota · Benchmark · CodeContestsHome/Leaderboards/Code & Software Engineering/Code Generation/CodeContests
Unknown

CodeContests.

13,610 competitive programming problems from CodeForces. ~200 private test cases per problem. 12+ programming languages.

Paper ↗Leaderboard ↓Lineage
§ 01 · Leaderboard

Results by metric.

Only 3 models on this benchmark
Help build the community leaderboard — submit your model results.
Found a wrong score or missing run?
Use row edits to send a sourced correction into moderation.
Add / edit result ↗Report issue ↗

pass@1

Pass@1 is the reported evaluation metric for CodeContests. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for pass@1verifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01GPT-4 + AlphaCodium
Fetched from CodeSOTA API on 2026-04-20
verified442026Source ↗Looks wrong?
02AlphaCode 2
Fetched from CodeSOTA API on 2026-04-20
verified432026Source ↗Looks wrong?
03GPT-4
Fetched from CodeSOTA API on 2026-04-20
verified192026Source ↗Looks wrong?
Lineage

CodeContests in context.

See full coding benchmarks lineage →
This benchmark (1)
active2022-02
CodeContests
None yet — this is the current frontier.
§ 04 · Submit a result

Add to the leaderboard.

Submit a Result

Sign in to submit benchmark results for CodeContests.

Sign in
← Back to Code Generation