Codesota · Benchmark · demon-benchHome/Leaderboards/demon-bench
Unknown

demon-bench.

demon-bench is a state-of-the-art machine learning benchmark indexed on Codesota. This page tracks published model results, top scores per metric, and the SOTA timeline for demon-bench.

Paper ↗Leaderboard ↓
§ 01 · Leaderboard

Results by metric.

Found a wrong score or missing run?
Use row edits to send a sourced correction into moderation.
Add / edit result ↗Report issue ↗

Multi Image Reasoning

Multi Image Reasoning is the reported evaluation metric for demon-bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Multi Image Reasoningverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Cheetah (Vicuna-13B)
Fetched from CodeSOTA API on 2026-04-20
verified53.652026Source ↗Looks wrong?
02Cheetah (Vicuna-7B)
Fetched from CodeSOTA API on 2026-04-20
verified50.282026Source ↗Looks wrong?
03Cheetah (LLaMA2-7B)
Fetched from CodeSOTA API on 2026-04-20
verified48.682026Source ↗Looks wrong?
04InstructBLIP
Fetched from CodeSOTA API on 2026-04-20
verified48.552026Source ↗Looks wrong?
05LLaMA-Adapter V2
Fetched from CodeSOTA API on 2026-04-20
verified44.032026Source ↗Looks wrong?
06Otter
Fetched from CodeSOTA API on 2026-04-20
verified43.852026Source ↗Looks wrong?
07MiniGPT-4
Fetched from CodeSOTA API on 2026-04-20
verified43.52026Source ↗Looks wrong?
08mPLUG-Owl
Fetched from CodeSOTA API on 2026-04-20
verified42.52026Source ↗Looks wrong?
09OpenFlamingo
Fetched from CodeSOTA API on 2026-04-20
verified41.632026Source ↗Looks wrong?
10LLaVA
Fetched from CodeSOTA API on 2026-04-20
verified41.532026Source ↗Looks wrong?
11BLIP-2
Fetched from CodeSOTA API on 2026-04-20
verified39.652026Source ↗Looks wrong?

Grounded Qa

Grounded Qa is the reported evaluation metric for demon-bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Grounded Qaverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Cheetah (Vicuna-13B)
Fetched from CodeSOTA API on 2026-04-20
verified52.932026Source ↗Looks wrong?
02Cheetah (LLaMA2-7B)
Fetched from CodeSOTA API on 2026-04-20
verified512026Source ↗Looks wrong?
03Cheetah (Vicuna-7B)
Fetched from CodeSOTA API on 2026-04-20
verified48.62026Source ↗Looks wrong?
04InstructBLIP
Fetched from CodeSOTA API on 2026-04-20
verified47.42026Source ↗Looks wrong?
05LLaMA-Adapter V2
Fetched from CodeSOTA API on 2026-04-20
verified44.82026Source ↗Looks wrong?
06Otter
Fetched from CodeSOTA API on 2026-04-20
verified41.672026Source ↗Looks wrong?
07BLIP-2
Fetched from CodeSOTA API on 2026-04-20
verified39.232026Source ↗Looks wrong?
08LLaVA
Fetched from CodeSOTA API on 2026-04-20
verified36.22026Source ↗Looks wrong?
09mPLUG-Owl
Fetched from CodeSOTA API on 2026-04-20
verified33.272026Source ↗Looks wrong?
10OpenFlamingo
Fetched from CodeSOTA API on 2026-04-20
verified322026Source ↗Looks wrong?
11MiniGPT-4
Fetched from CodeSOTA API on 2026-04-20
verified30.272026Source ↗Looks wrong?

Knowledge Images Qa

Knowledge Images Qa is the reported evaluation metric for demon-bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Knowledge Images Qaverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Cheetah (Vicuna-13B)
Fetched from CodeSOTA API on 2026-04-20
verified49.332026Source ↗Looks wrong?
02Cheetah (Vicuna-7B)
Fetched from CodeSOTA API on 2026-04-20
verified44.932026Source ↗Looks wrong?
03Cheetah (LLaMA2-7B)
Fetched from CodeSOTA API on 2026-04-20
verified44.932026Source ↗Looks wrong?
04InstructBLIP
Fetched from CodeSOTA API on 2026-04-20
verified44.42026Source ↗Looks wrong?
05BLIP-2
Fetched from CodeSOTA API on 2026-04-20
verified33.532026Source ↗Looks wrong?
06mPLUG-Owl
Fetched from CodeSOTA API on 2026-04-20
verified32.472026Source ↗Looks wrong?
07LLaMA-Adapter V2
Fetched from CodeSOTA API on 2026-04-20
verified322026Source ↗Looks wrong?
08OpenFlamingo
Fetched from CodeSOTA API on 2026-04-20
verified30.62026Source ↗Looks wrong?
09LLaVA
Fetched from CodeSOTA API on 2026-04-20
verified28.332026Source ↗Looks wrong?
10Otter
Fetched from CodeSOTA API on 2026-04-20
verified27.732026Source ↗Looks wrong?
11MiniGPT-4
Fetched from CodeSOTA API on 2026-04-20
verified26.42026Source ↗Looks wrong?

Multimodal Dialogue

Multimodal Dialogue is the reported evaluation metric for demon-bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Multimodal Dialogueverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Cheetah (LLaMA2-7B)
Fetched from CodeSOTA API on 2026-04-20
verified42.72026Source ↗Looks wrong?
02Cheetah (Vicuna-13B)
Fetched from CodeSOTA API on 2026-04-20
verified38.142026Source ↗Looks wrong?
03Cheetah (Vicuna-7B)
Fetched from CodeSOTA API on 2026-04-20
verified37.52026Source ↗Looks wrong?
04InstructBLIP
Fetched from CodeSOTA API on 2026-04-20
verified33.582026Source ↗Looks wrong?
05BLIP-2
Fetched from CodeSOTA API on 2026-04-20
verified26.122026Source ↗Looks wrong?
06OpenFlamingo
Fetched from CodeSOTA API on 2026-04-20
verified16.882026Source ↗Looks wrong?
07Otter
Fetched from CodeSOTA API on 2026-04-20
verified15.372026Source ↗Looks wrong?
08LLaMA-Adapter V2
Fetched from CodeSOTA API on 2026-04-20
verified14.222026Source ↗Looks wrong?
09MiniGPT-4
Fetched from CodeSOTA API on 2026-04-20
verified13.692026Source ↗Looks wrong?
10mPLUG-Owl
Fetched from CodeSOTA API on 2026-04-20
verified12.672026Source ↗Looks wrong?
11LLaVA
Fetched from CodeSOTA API on 2026-04-20
verified7.792026Source ↗Looks wrong?

Accuracy

Accuracy is the reported evaluation metric for demon-bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Accuracyverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Cheetah (Vicuna-13B)
Fetched from CodeSOTA API on 2026-04-20
verified39.282026Source ↗Looks wrong?
02Cheetah (LLaMA2-7B)
Fetched from CodeSOTA API on 2026-04-20
verified37.222026Source ↗Looks wrong?
03Cheetah (Vicuna-7B)
Fetched from CodeSOTA API on 2026-04-20
verified36.372026Source ↗Looks wrong?
04InstructBLIP
Fetched from CodeSOTA API on 2026-04-20
verified332026Source ↗Looks wrong?
05BLIP-2
Fetched from CodeSOTA API on 2026-04-20
verified26.922026Source ↗Looks wrong?
06LLaMA-Adapter V2
Fetched from CodeSOTA API on 2026-04-20
verified26.32026Source ↗Looks wrong?
07OpenFlamingo
Fetched from CodeSOTA API on 2026-04-20
verified25.832026Source ↗Looks wrong?
08Otter
Fetched from CodeSOTA API on 2026-04-20
verified24.512026Source ↗Looks wrong?
09mPLUG-Owl
Fetched from CodeSOTA API on 2026-04-20
verified23.132026Source ↗Looks wrong?
10MiniGPT-4
Fetched from CodeSOTA API on 2026-04-20
verified22.212026Source ↗Looks wrong?
11LLaVA
Fetched from CodeSOTA API on 2026-04-20
verified21.242026Source ↗Looks wrong?

Visual Inference

Visual Inference is the reported evaluation metric for demon-bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Visual Inferenceverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Cheetah (Vicuna-13B)
Fetched from CodeSOTA API on 2026-04-20
verified27.152026Source ↗Looks wrong?
02Cheetah (Vicuna-7B)
Fetched from CodeSOTA API on 2026-04-20
verified25.92026Source ↗Looks wrong?
03Cheetah (LLaMA2-7B)
Fetched from CodeSOTA API on 2026-04-20
verified25.52026Source ↗Looks wrong?
04OpenFlamingo
Fetched from CodeSOTA API on 2026-04-20
verified13.852026Source ↗Looks wrong?
05LLaMA-Adapter V2
Fetched from CodeSOTA API on 2026-04-20
verified13.512026Source ↗Looks wrong?
06InstructBLIP
Fetched from CodeSOTA API on 2026-04-20
verified11.492026Source ↗Looks wrong?
07Otter
Fetched from CodeSOTA API on 2026-04-20
verified11.392026Source ↗Looks wrong?
08BLIP-2
Fetched from CodeSOTA API on 2026-04-20
verified10.672026Source ↗Looks wrong?
09LLaVA
Fetched from CodeSOTA API on 2026-04-20
verified8.272026Source ↗Looks wrong?
10MiniGPT-4
Fetched from CodeSOTA API on 2026-04-20
verified7.952026Source ↗Looks wrong?
11mPLUG-Owl
Fetched from CodeSOTA API on 2026-04-20
verified5.402026Source ↗Looks wrong?

Relation Cloze

Relation Cloze is the reported evaluation metric for demon-bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Relation Clozeverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Cheetah (Vicuna-13B)
Fetched from CodeSOTA API on 2026-04-20
verified27.152026Source ↗Looks wrong?
02Cheetah (LLaMA2-7B)
Fetched from CodeSOTA API on 2026-04-20
verified22.952026Source ↗Looks wrong?
03Cheetah (Vicuna-7B)
Fetched from CodeSOTA API on 2026-04-20
verified22.152026Source ↗Looks wrong?
04OpenFlamingo
Fetched from CodeSOTA API on 2026-04-20
verified21.652026Source ↗Looks wrong?
05InstructBLIP
Fetched from CodeSOTA API on 2026-04-20
verified21.22026Source ↗Looks wrong?
06LLaMA-Adapter V2
Fetched from CodeSOTA API on 2026-04-20
verified182026Source ↗Looks wrong?
07BLIP-2
Fetched from CodeSOTA API on 2026-04-20
verified17.942026Source ↗Looks wrong?
08MiniGPT-4
Fetched from CodeSOTA API on 2026-04-20
verified16.62026Source ↗Looks wrong?
09mPLUG-Owl
Fetched from CodeSOTA API on 2026-04-20
verified16.252026Source ↗Looks wrong?
10Otter
Fetched from CodeSOTA API on 2026-04-20
verified162026Source ↗Looks wrong?
11LLaVA
Fetched from CodeSOTA API on 2026-04-20
verified15.852026Source ↗Looks wrong?

Storytelling

Storytelling is the reported evaluation metric for demon-bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Storytellingverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Cheetah (Vicuna-13B)
Fetched from CodeSOTA API on 2026-04-20
verified26.592026Source ↗Looks wrong?
02Cheetah (Vicuna-7B)
Fetched from CodeSOTA API on 2026-04-20
verified25.22026Source ↗Looks wrong?
03Cheetah (LLaMA2-7B)
Fetched from CodeSOTA API on 2026-04-20
verified24.762026Source ↗Looks wrong?
04InstructBLIP
Fetched from CodeSOTA API on 2026-04-20
verified24.412026Source ↗Looks wrong?
05OpenFlamingo
Fetched from CodeSOTA API on 2026-04-20
verified24.222026Source ↗Looks wrong?
06BLIP-2
Fetched from CodeSOTA API on 2026-04-20
verified21.312026Source ↗Looks wrong?
07mPLUG-Owl
Fetched from CodeSOTA API on 2026-04-20
verified19.332026Source ↗Looks wrong?
08LLaMA-Adapter V2
Fetched from CodeSOTA API on 2026-04-20
verified17.572026Source ↗Looks wrong?
09MiniGPT-4
Fetched from CodeSOTA API on 2026-04-20
verified17.072026Source ↗Looks wrong?
10Otter
Fetched from CodeSOTA API on 2026-04-20
verified15.572026Source ↗Looks wrong?
11LLaVA
Fetched from CodeSOTA API on 2026-04-20
verified10.72026Source ↗Looks wrong?
§ 04 · Submit a result

Add to the leaderboard.

Submit a Result

Sign in to submit benchmark results for demon-bench.

Sign in
← Back to Leaderboards