Browse State-of-the-Art › Environmental Sound Classification
Environmental Sound Classification
27 papers with code · 3 benchmarks · 6 datasets archive 2025-07-28
Classification of Environmental Sounds. Most often sounds found in Urban environments. Task related to noise monitoring.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| UrbanSound8K (3 rows) | AudioCLIP | AudioCLIP: Extending CLIP to Image, Text and Audio | code | Syntology ran 0 of 6 samples · 6 unverified | Compare |
| ESC-50 (1 row) | AudioCLIP | AudioCLIP: Extending CLIP to Image, Text and Audio | code | Syntology ran 0 of 6 samples · 6 unverified | Compare |
| FSD50K (1 row) | [ABT] AudioNTT | Audio Barlow Twins: Self-Supervised Audio Representation Learning | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
6 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
27 shown of 27 papers with code (46 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
15 Aug 2016 5 repositories listedWe show that the improved performance stems from the combination of a deep, high-capacity model and an augmented training set: this combination outperforms both the proposed CNN without augmentation and a "shallow"…
-
24 Jun 2021 4 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedAudioCLIP achieves new state-of-the-art results in the Environmental Sound Classification (ESC) task, out-performing other approaches by reaching accuracies of 90.
-
22 Jul 2020 3 repositories listedBesides, we show that even though we use the pretrained model weights for initialization, there is variance in performance in various output runs of the same model.
-
20 May 2022 2 repositories listedPaddleSpeech is an open-source all-in-one speech toolkit.
-
24 Aug 2020 2 repositories listedThis paper describes CRNNs we used to participate in Task 5 of the DCASE 2020 challenge.
-
18 Apr 2019 2 repositories listedIn this paper, we present an end-to-end approach for environmental sound classification based on a 1D Convolution Neural Network (CNN) that learns a representation directly from the audio signal.
-
25 May 2018 2 repositories listedWe have evaluated the MCLNN performance using the Urbansound8k dataset of environmental sounds.
-
4 Jun 2025 1 repository listedAudio-text models are widely used in zero-shot environmental sound classification as they alleviate the need for annotated data.
-
11 Mar 2025 1 repository listedConvolutional Dictionary Learning (CDL) has emerged as a powerful approach for signal representation by learning translation-invariant features through convolution operations.
-
17 Feb 2025 1 repository listedTherefore, we propose a new method: MAsked latenT Prediction And Classification (MATPAC), which is trained with two pretext tasks solved jointly.
-
7 Mar 2023 1 repository listedThis paper presents a context-aware framework for feature selection and classification procedures to realize a fast and accurate audio event annotation and classification.
-
Effective Audio Classification Network Based on Paired Inverse Pyramid Structure and Dense MLP Block5 Nov 2022 1 repository listedRecently, massive architectures based on Convolutional Neural Network (CNN) and self-attention mechanisms have become necessary for audio classification.
-
28 Sep 2022 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)The Barlow Twins self-supervised learning objective requires neither negative samples or asymmetric learning updates, achieving results on a par with the current state-of-the-art within Computer Vision.
-
15 Jul 2022 1 repository listedExperimental results on the DCASE 2019 Task 1 and ESC-50 dataset show that our proposed method outperforms baseline continual learning methods on classification accuracy and computational efficiency, indicating our…
-
25 Apr 2022 1 repository listedWhile efficient architectures and a plethora of augmentations for end-to-end image classification tasks have been suggested and heavily investigated, state-of-the-art techniques for audio classifications still rely on…
-
15 Mar 2022 1 repository listedAUCO ResNet has proved to provide state of art results on many datasets.
-
22 Mar 2021 1 repository listedWith the growth of the Internet of Things and the rise of Big Data, data processing and machine learning applications are being moved to cheap and low size, weight, and power (SWaP) devices at the edge, often in the…
-
5 Mar 2021 1 repository listedSignificant efforts are being invested to bring state-of-the-art classification and recognition to edge devices with extreme resource constraints (memory, speed, and lack of GPU support).
-
2 Mar 2021 1 repository listedOur extensive benchmark experiments show that our hybrid deep network models trained with combined contrastive and cross-entropy loss achieved the state-of-the-art performance on three benchmark datasets ESC-10, ESC-50,…
-
16 Feb 2021 1 repository listedIn all but one cases, MM, RMM, and FM outperformed MT and DCT significantly, MM and RMM being the best methods in most experiments.
-
22 Oct 2020 1 repository listedSometimes authors copy-pasting the results of the original papers which is not helping reproducibility.
-
15 Apr 2020 1 repository listedEnvironmental Sound Classification (ESC) is an active research area in the audio domain and has seen a lot of progress in the past years.
-
27 Sep 2019 1 repository listedThe proposed model uses log-scaled Mel-spectrogram as the representation format for the audio data.
-
15 May 2019 1 repository listedNoise monitoring using Wireless Sensor Networks are being applied in order to understand and help mitigate these noise problems.
-
14 Oct 2018 1 repository listedDespite sound being a rich source of information, computing devices with microphones do not leverage audio to glean useful insights about their physical and social context.
-
1 Dec 2017 1 repository listedEnd-to-end neural network based approaches to audio modelling are generally outperformed by models trained on high-level data representations.
-
23 Jan 2017 1 repository listedWe show that the improved performance stems from the combination of a deep, high-capacity model and an augmented training set: this combination outperforms both the proposed CNN without augmentation and a “shallow”…
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections