Codesota · Benchmark · audiocapsHome/Leaderboards/audiocaps
Unknown

audiocaps.

audiocaps is a state-of-the-art machine learning benchmark indexed on Codesota. This page tracks published model results, top scores per metric, and the SOTA timeline for audiocaps.

Paper ↗Leaderboard ↓
§ 01 · Leaderboard

Results by metric.

Found a wrong score or missing run?
Use row edits to send a sourced correction into moderation.
Add / edit result ↗Report issue ↗

Fad

Fad is the reported evaluation metric for audiocaps. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Fadverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01AudioLDM
Fetched from CodeSOTA API on 2026-04-20
verified4.482026Source ↗Looks wrong?
02AudioLDM 2-Full-Large
Fetched from CodeSOTA API on 2026-04-20
verified1.862026Source ↗Looks wrong?
03AudioLDM 2-Full
Fetched from CodeSOTA API on 2026-04-20
verified1.782026Source ↗Looks wrong?
04TANGO
Fetched from CodeSOTA API on 2026-04-20
verified1.732026Source ↗Looks wrong?
05AudioLDM 2-AC-Large
Fetched from CodeSOTA API on 2026-04-20
verified1.422026Source ↗Looks wrong?
§ 04 · Submit a result

Add to the leaderboard.

Submit a Result

Sign in to submit benchmark results for audiocaps.

Sign in
← Back to Leaderboards