Codesota · Benchmark · ESC-50Home/Leaderboards/Audio & Speech/Audio Classification/ESC-50
Unknown

ESC-50.

2,000 environmental audio recordings organized into 50 classes (animals, natural soundscapes, etc.).

Paper ↗Leaderboard ↓Lineage
§ 01 · Leaderboard

Results by metric.

Only 4 models on this benchmark
Help build the community leaderboard — submit your model results.
Found a wrong score or missing run?
Use row edits to send a sourced correction into moderation.
Add / edit result ↗Report issue ↗

accuracy

Accuracy is the reported evaluation metric for ESC-50. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for accuracyverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01BEATs
Fetched from CodeSOTA API on 2026-04-20
verified98.12026Source ↗Looks wrong?
02HTS-AT
Fetched from CodeSOTA API on 2026-04-20
verified972026Source ↗Looks wrong?
03AST
Fetched from CodeSOTA API on 2026-04-20
verified95.62026Source ↗Looks wrong?
04CLAP
Fetched from CodeSOTA API on 2026-04-20
verified93.72026Source ↗Looks wrong?
Lineage

ESC-50 in context.

See full audio understanding benchmarks lineage →
None — this is where the lineage begins.
This benchmark (1)
saturated2015-01
ESC-50
Successors (2)
saturating2017-03
AudioSet
AudioSet replaced ESC-50 as the primary audio classification benchmark — 527 classes vs 50, 2M clips vs 2K, hierarchical ontology. Scale and coverage made it the ImageNet analogue for audio. ESC-50 became a probe task for pretrained representations.
active2017-10
MUSDB18
MUSDB18 branches into music source separation — a generative audio task, not a classification one. Different task family entirely; ESC-50's sound-class framework doesn't apply.
§ 04 · Submit a result

Add to the leaderboard.

Submit a Result

Sign in to submit benchmark results for ESC-50.

Sign in
← Back to Audio Classification