Codesota · Models · Grok-4.20xAI7 results · 1 benchmarks
Model card

Grok-4.20.

xAIopen-source
§ 02 · Benchmarks

Every benchmark Grok-4.20 has a recorded score for.

#BenchmarkArea · TaskMetricValueRankDateSource
01PLCCNatural Language Processing · Polish Cultural Competencygrammar72.0%#42/165—source ↗
02PLCCNatural Language Processing · Polish Cultural Competencyhistory82.0%#53/165—source ↗
03PLCCNatural Language Processing · Polish Cultural Competencyaverage67.8%#71/165—source ↗
04PLCCNatural Language Processing · Polish Cultural Competencyvocabulary59.0%#75/165—source ↗
05PLCCNatural Language Processing · Polish Cultural Competencyculture-and-tradition65.0%#76/165—source ↗
06PLCCNatural Language Processing · Polish Cultural Competencyart-and-entertainment55.0%#77/165—source ↗
07PLCCNatural Language Processing · Polish Cultural Competencygeography74.0%#80/165—source ↗
Rank column shows this model’s position vs all other models scored on the same benchmark + metric (competitors after the slash). #1 in red means current SOTA. Sorted by rank, then newest result.
§ 03 · Strengths by area

Where Grok-4.20 actually performs.

Natural Language Processing
1
benchmark
avg rank #67.7
§ 05 · Related models

Other xAI models scored on Codesota.

Grok 4
15 results
Grok-2-1212
7 results
Grok-3-Beta
7 results
Grok-3-Mini-Beta
7 results
Grok-4.1-Fast
7 results
Grok-4-Fast
7 results
Grok 2
4 results
Grok 3
1 result
§ 06 · Sources & freshness

Where these numbers come from.

sdadas/PLCC
7
results
7 of 7 rows marked verified.