Codesota · Benchmark · AudioSetHome/Leaderboards/Audio & Speech/Audio Classification/AudioSet
Unknown

AudioSet.

2M+ human-labeled 10-second YouTube video clips covering 632 audio event classes.

Paper ↗Leaderboard ↓Lineage
§ 01 · Leaderboard

Results by metric.

Only 4 models on this benchmark
Help build the community leaderboard — submit your model results.
Found a wrong score or missing run?
Use row edits to send a sourced correction into moderation.
Add / edit result ↗Report issue ↗

map

Map is the reported evaluation metric for AudioSet. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for mapverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01BEATs
Fetched from CodeSOTA API on 2026-04-20
verified0.512026Source ↗Looks wrong?
02AST
Fetched from CodeSOTA API on 2026-04-20
verified0.482026Source ↗Looks wrong?
03HTS-AT
Fetched from CodeSOTA API on 2026-04-20
verified0.472026Source ↗Looks wrong?
04CLAP
Fetched from CodeSOTA API on 2026-04-20
verified0.432026Source ↗Looks wrong?
Lineage

AudioSet in context.

See full audio understanding benchmarks lineage →
This benchmark (1)
saturating2017-03
AudioSet
Successors (1)
active2020-01
Clotho
Clotho shifted the evaluation task from classification (what sounds are here?) to captioning (describe these sounds in a sentence). A scope shift enabled by the growing capability of audio encoders trained on AudioSet.
§ 04 · Submit a result

Add to the leaderboard.

Submit a Result

Sign in to submit benchmark results for AudioSet.

Sign in
← Back to Audio Classification