demon-bench is a state-of-the-art machine learning benchmark indexed on Codesota. This page tracks published model results, top scores per metric, and the SOTA timeline for demon-bench.
Multi Image Reasoning is the reported evaluation metric for demon-bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Cheetah (Vicuna-13B) | verified | 53.65 | 2026 | Source ↗ | Looks wrong? |
| 02 | Cheetah (Vicuna-7B) | verified | 50.28 | 2026 | Source ↗ | Looks wrong? |
| 03 | Cheetah (LLaMA2-7B) | verified | 48.68 | 2026 | Source ↗ | Looks wrong? |
| 04 | InstructBLIP | verified | 48.55 | 2026 | Source ↗ | Looks wrong? |
| 05 | LLaMA-Adapter V2 | verified | 44.03 | 2026 | Source ↗ | Looks wrong? |
| 06 | Otter | verified | 43.85 | 2026 | Source ↗ | Looks wrong? |
| 07 | MiniGPT-4 | verified | 43.5 | 2026 | Source ↗ | Looks wrong? |
| 08 | mPLUG-Owl | verified | 42.5 | 2026 | Source ↗ | Looks wrong? |
| 09 | OpenFlamingo | verified | 41.63 | 2026 | Source ↗ | Looks wrong? |
| 10 | LLaVA | verified | 41.53 | 2026 | Source ↗ | Looks wrong? |
| 11 | BLIP-2 | verified | 39.65 | 2026 | Source ↗ | Looks wrong? |
Grounded Qa is the reported evaluation metric for demon-bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Cheetah (Vicuna-13B) | verified | 52.93 | 2026 | Source ↗ | Looks wrong? |
| 02 | Cheetah (LLaMA2-7B) | verified | 51 | 2026 | Source ↗ | Looks wrong? |
| 03 | Cheetah (Vicuna-7B) | verified | 48.6 | 2026 | Source ↗ | Looks wrong? |
| 04 | InstructBLIP | verified | 47.4 | 2026 | Source ↗ | Looks wrong? |
| 05 | LLaMA-Adapter V2 | verified | 44.8 | 2026 | Source ↗ | Looks wrong? |
| 06 | Otter | verified | 41.67 | 2026 | Source ↗ | Looks wrong? |
| 07 | BLIP-2 | verified | 39.23 | 2026 | Source ↗ | Looks wrong? |
| 08 | LLaVA | verified | 36.2 | 2026 | Source ↗ | Looks wrong? |
| 09 | mPLUG-Owl | verified | 33.27 | 2026 | Source ↗ | Looks wrong? |
| 10 | OpenFlamingo | verified | 32 | 2026 | Source ↗ | Looks wrong? |
| 11 | MiniGPT-4 | verified | 30.27 | 2026 | Source ↗ | Looks wrong? |
Knowledge Images Qa is the reported evaluation metric for demon-bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Cheetah (Vicuna-13B) | verified | 49.33 | 2026 | Source ↗ | Looks wrong? |
| 02 | Cheetah (Vicuna-7B) | verified | 44.93 | 2026 | Source ↗ | Looks wrong? |
| 03 | Cheetah (LLaMA2-7B) | verified | 44.93 | 2026 | Source ↗ | Looks wrong? |
| 04 | InstructBLIP | verified | 44.4 | 2026 | Source ↗ | Looks wrong? |
| 05 | BLIP-2 | verified | 33.53 | 2026 | Source ↗ | Looks wrong? |
| 06 | mPLUG-Owl | verified | 32.47 | 2026 | Source ↗ | Looks wrong? |
| 07 | LLaMA-Adapter V2 | verified | 32 | 2026 | Source ↗ | Looks wrong? |
| 08 | OpenFlamingo | verified | 30.6 | 2026 | Source ↗ | Looks wrong? |
| 09 | LLaVA | verified | 28.33 | 2026 | Source ↗ | Looks wrong? |
| 10 | Otter | verified | 27.73 | 2026 | Source ↗ | Looks wrong? |
| 11 | MiniGPT-4 | verified | 26.4 | 2026 | Source ↗ | Looks wrong? |
Multimodal Dialogue is the reported evaluation metric for demon-bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Cheetah (LLaMA2-7B) | verified | 42.7 | 2026 | Source ↗ | Looks wrong? |
| 02 | Cheetah (Vicuna-13B) | verified | 38.14 | 2026 | Source ↗ | Looks wrong? |
| 03 | Cheetah (Vicuna-7B) | verified | 37.5 | 2026 | Source ↗ | Looks wrong? |
| 04 | InstructBLIP | verified | 33.58 | 2026 | Source ↗ | Looks wrong? |
| 05 | BLIP-2 | verified | 26.12 | 2026 | Source ↗ | Looks wrong? |
| 06 | OpenFlamingo | verified | 16.88 | 2026 | Source ↗ | Looks wrong? |
| 07 | Otter | verified | 15.37 | 2026 | Source ↗ | Looks wrong? |
| 08 | LLaMA-Adapter V2 | verified | 14.22 | 2026 | Source ↗ | Looks wrong? |
| 09 | MiniGPT-4 | verified | 13.69 | 2026 | Source ↗ | Looks wrong? |
| 10 | mPLUG-Owl | verified | 12.67 | 2026 | Source ↗ | Looks wrong? |
| 11 | LLaVA | verified | 7.79 | 2026 | Source ↗ | Looks wrong? |
Accuracy is the reported evaluation metric for demon-bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Cheetah (Vicuna-13B) | verified | 39.28 | 2026 | Source ↗ | Looks wrong? |
| 02 | Cheetah (LLaMA2-7B) | verified | 37.22 | 2026 | Source ↗ | Looks wrong? |
| 03 | Cheetah (Vicuna-7B) | verified | 36.37 | 2026 | Source ↗ | Looks wrong? |
| 04 | InstructBLIP | verified | 33 | 2026 | Source ↗ | Looks wrong? |
| 05 | BLIP-2 | verified | 26.92 | 2026 | Source ↗ | Looks wrong? |
| 06 | LLaMA-Adapter V2 | verified | 26.3 | 2026 | Source ↗ | Looks wrong? |
| 07 | OpenFlamingo | verified | 25.83 | 2026 | Source ↗ | Looks wrong? |
| 08 | Otter | verified | 24.51 | 2026 | Source ↗ | Looks wrong? |
| 09 | mPLUG-Owl | verified | 23.13 | 2026 | Source ↗ | Looks wrong? |
| 10 | MiniGPT-4 | verified | 22.21 | 2026 | Source ↗ | Looks wrong? |
| 11 | LLaVA | verified | 21.24 | 2026 | Source ↗ | Looks wrong? |
Visual Inference is the reported evaluation metric for demon-bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Cheetah (Vicuna-13B) | verified | 27.15 | 2026 | Source ↗ | Looks wrong? |
| 02 | Cheetah (Vicuna-7B) | verified | 25.9 | 2026 | Source ↗ | Looks wrong? |
| 03 | Cheetah (LLaMA2-7B) | verified | 25.5 | 2026 | Source ↗ | Looks wrong? |
| 04 | OpenFlamingo | verified | 13.85 | 2026 | Source ↗ | Looks wrong? |
| 05 | LLaMA-Adapter V2 | verified | 13.51 | 2026 | Source ↗ | Looks wrong? |
| 06 | InstructBLIP | verified | 11.49 | 2026 | Source ↗ | Looks wrong? |
| 07 | Otter | verified | 11.39 | 2026 | Source ↗ | Looks wrong? |
| 08 | BLIP-2 | verified | 10.67 | 2026 | Source ↗ | Looks wrong? |
| 09 | LLaVA | verified | 8.27 | 2026 | Source ↗ | Looks wrong? |
| 10 | MiniGPT-4 | verified | 7.95 | 2026 | Source ↗ | Looks wrong? |
| 11 | mPLUG-Owl | verified | 5.40 | 2026 | Source ↗ | Looks wrong? |
Relation Cloze is the reported evaluation metric for demon-bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Cheetah (Vicuna-13B) | verified | 27.15 | 2026 | Source ↗ | Looks wrong? |
| 02 | Cheetah (LLaMA2-7B) | verified | 22.95 | 2026 | Source ↗ | Looks wrong? |
| 03 | Cheetah (Vicuna-7B) | verified | 22.15 | 2026 | Source ↗ | Looks wrong? |
| 04 | OpenFlamingo | verified | 21.65 | 2026 | Source ↗ | Looks wrong? |
| 05 | InstructBLIP | verified | 21.2 | 2026 | Source ↗ | Looks wrong? |
| 06 | LLaMA-Adapter V2 | verified | 18 | 2026 | Source ↗ | Looks wrong? |
| 07 | BLIP-2 | verified | 17.94 | 2026 | Source ↗ | Looks wrong? |
| 08 | MiniGPT-4 | verified | 16.6 | 2026 | Source ↗ | Looks wrong? |
| 09 | mPLUG-Owl | verified | 16.25 | 2026 | Source ↗ | Looks wrong? |
| 10 | Otter | verified | 16 | 2026 | Source ↗ | Looks wrong? |
| 11 | LLaVA | verified | 15.85 | 2026 | Source ↗ | Looks wrong? |
Storytelling is the reported evaluation metric for demon-bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Cheetah (Vicuna-13B) | verified | 26.59 | 2026 | Source ↗ | Looks wrong? |
| 02 | Cheetah (Vicuna-7B) | verified | 25.2 | 2026 | Source ↗ | Looks wrong? |
| 03 | Cheetah (LLaMA2-7B) | verified | 24.76 | 2026 | Source ↗ | Looks wrong? |
| 04 | InstructBLIP | verified | 24.41 | 2026 | Source ↗ | Looks wrong? |
| 05 | OpenFlamingo | verified | 24.22 | 2026 | Source ↗ | Looks wrong? |
| 06 | BLIP-2 | verified | 21.31 | 2026 | Source ↗ | Looks wrong? |
| 07 | mPLUG-Owl | verified | 19.33 | 2026 | Source ↗ | Looks wrong? |
| 08 | LLaMA-Adapter V2 | verified | 17.57 | 2026 | Source ↗ | Looks wrong? |
| 09 | MiniGPT-4 | verified | 17.07 | 2026 | Source ↗ | Looks wrong? |
| 10 | Otter | verified | 15.57 | 2026 | Source ↗ | Looks wrong? |
| 11 | LLaVA | verified | 10.7 | 2026 | Source ↗ | Looks wrong? |