Datasets › ESC-50

ESC-50

Introduced in ESC: Dataset for Environmental Sound Classification1 Jan 2015 archive 2025-07-28

The ESC-50 dataset is a labeled collection of 2000 environmental audio recordings suitable for benchmarking methods of environmental sound classification. It comprises 2000 5s-clips of 50 different classes across natural, human and domestic sounds, again, drawn from Freesound.org.

Source: The NIGENS General Sound Events Database Image Source: https://github.com/karolpiczak/ESC-50

Benchmarks archive 2025-07-28

All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

27 shown of 27 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 387. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP 2 1 28 Mar 2025 not harvested
Masked Latent Prediction and Classification for Self-Supervised Audio Representation Learning 1 1 17 Feb 2025 not harvested
LHGNN: Local-Higher Order Graph Neural Networks For Audio Classification and Tagging 0 1 7 Jan 2025 not harvested
Performance of Gaussian Mixture Model Classifiers on Embedded Feature Spaces 1 1 17 Oct 2024 not harvested
M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation 2 1 4 Jun 2024 not harvested
Masked Modeling Duo: Towards a Universal Audio Pre-training Framework 2 2 9 Apr 2024 ran 7 of 7 samples (0 unverified; 7 pointer-only for licence)
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding 2 1 22 Mar 2024 not harvested
EAT: Self-Supervised Pre-Training with Efficient Audio Transformer 1 1 7 Jan 2024 ran 13 of 17 samples (4 unverified)
OmniVec2 - A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning 0 1 1 Jan 2024 not harvested
OmniVec: Learning robust representations with cross modal sharing 0 1 7 Nov 2023 not harvested
Dynamic Convolutional Neural Networks as Efficient Pre-trained Audio Models 1 1 24 Oct 2023 not harvested
MT-SLVR: Multi-Task Self-Supervised Learning for Transformation In(Variant) Representations 1 3 29 May 2023 not harvested
BEATs: Audio Pre-Training with Acoustic Tokenizers 4 1 18 Dec 2022 ran 3 of 18 samples (15 unverified)
Efficient Large-scale Audio Tagging via Transformer-to-CNN Knowledge Distillation 2 1 9 Nov 2022 ran 1 of 2 samples (1 unverified)
Learning Rate Curriculum 1 1 18 May 2022 not harvested
End-to-End Audio Strikes Back: Boosting Augmentations Towards An Efficient Audio Classification Network 1 3 25 Apr 2022 not harvested
MetaAudio: A Few-Shot Audio Classification Benchmark 1 7 5 Apr 2022 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
SepTr: Separable Transformer for Audio Spectrogram Processing 1 1 17 Mar 2022 not harvested
HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection 1 1 2 Feb 2022 ran 4 of 9 samples (5 unverified)
AudioCLIP: Extending CLIP to Image, Text and Audio 4 1 24 Jun 2021 ran 0 of 6 samples (6 unverified)
AST: Audio Spectrogram Transformer 5 1 5 Apr 2021 ran 0 of 3 samples (3 unverified)
Multi-Format Contrastive Learning of Audio Representations 0 1 11 Mar 2021 not harvested
Environmental Sound Classification on the Edge: A Pipeline for Deep Acoustic Networks on Extremely Resource-Constrained Devices 1 1 5 Mar 2021 not harvested
Audio-Visual Instance Discrimination with Cross-Modal Agreement 1 1 27 Apr 2020 not harvested
Self-Supervised Learning by Cross-Modal Audio-Video Clustering 1 2 28 Nov 2019 ran 0 of 1 samples (1 unverified)
Cooperative Learning of Audio and Video Models from Self-Supervised Synchronization 0 1 30 Jun 2018 not harvested
Look, Listen and Learn 1 1 23 May 2017 not harvested

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY-NC 3.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • ESC-50

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections