Codesota · Benchmark · NoCapsHome/Leaderboards/Multimodal Media/Image Captioning/NoCaps
Unknown

NoCaps.

15K validation images from Open Images with 166K human-written captions. Specifically tests zero-shot generalization to novel objects not seen during training.

Paper ↗Leaderboard ↓
§ 01 · Leaderboard

Results by metric.

Found a wrong score or missing run?
Use row edits to send a sourced correction into moderation.
Add / edit result ↗Report issue ↗

cider

Cider is the reported evaluation metric for NoCaps. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for ciderverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01CogVLM-17B
Fetched from CodeSOTA API on 2026-04-20
verified128.32026Source ↗Looks wrong?
02PaLI-X-55B
Fetched from CodeSOTA API on 2026-04-20
verified126.32026Source ↗Looks wrong?
03PaLI-17B
Fetched from CodeSOTA API on 2026-04-20
verified124.42026Source ↗Looks wrong?
04BLIP-2 (FlanT5XL)
Fetched from CodeSOTA API on 2026-04-20
verified123.72026Source ↗Looks wrong?
05BLIP-2 (OPT 2.7B)
Fetched from CodeSOTA API on 2026-04-20
verified121.62026Source ↗Looks wrong?
§ 04 · Submit a result

Add to the leaderboard.

Submit a Result

Sign in to submit benchmark results for NoCaps.

Sign in
← Back to Image Captioning