Browse State-of-the-Art › Acoustic Scene Classification
Acoustic Scene Classification
42 papers with code · 5 benchmarks · 10 datasets archive 2025-07-28
The goal of acoustic scene classification is to classify a test recording into one of the provided predefined classes that characterizes the environment in which it was recorded.
Source: DCASE 2019 Source: DCASE 2018
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
5 leaderboard tables shown for this task, 5 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| CochlScene (2 rows) | Audio Flamingo | Audio Flamingo: A Novel Audio Language Model with Few-Shot... | code | Syntology ran 3 of 3 samples · 0 unverified | Compare |
| DCASE 2019 Mobile (1 row) | Basic + Spectrum Correction | Spectrum Correction: Acoustic Scene Classification with Mismatched... | code | — | Compare |
| TAU Urban Acoustic Scenes 2019 (1 row) | Two-stage ensemble system | A Two-Stage Approach to Device-Robust Acoustic Scene Classification | code | Syntology ran 0 of 10 samples · 10 unverified | Compare |
| TUT Acoustic Scenes 2017 (1 row) | Qwen-Audio | Qwen-Audio: Advancing Universal Audio Understanding via Unified... | code | Syntology ran 5 of 7 samples · 2 unverified | Compare |
| TUT Urban Acoustic Scenes 2018 (1 row) | ERGL: event relational graph representation learning | Multi-dimensional Edge-based Audio Event Relational Graph... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
10 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 42 papers with code (132 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
3 Jul 2019 3 repositories listedTo this end, we analyse the receptive field (RF) of these CNNs and demonstrate the importance of the RF to the generalization capability of the models.
-
14 Nov 2023 2 repositories listed Syntology ran 5 of 7 samples · 2 unverified · 7 pointer-only (licence)Recently, instruction-following audio-language models have received broad attention for audio interaction with humans.
-
11 Oct 2021 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedHowever, one of the main shortcomings of transformer models, compared to the well-established CNNs, is the computational complexity.
-
3 Mar 2020 2 repositories listedThe understanding of the surrounding environment plays a critical role in autonomous robotic systems, such as self-driving cars.
-
5 Sep 2019 2 repositories listedOne side effect of restricting the RF of CNNs is that more frequency information is lost.
-
24 Oct 2018 2 repositories listedWe investigate supervised learning strategies that improve the training of neural network audio classifiers on small annotated collections.
-
25 Jul 2018 2 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedThis paper introduces the acoustic scene classification task of DCASE 2018 Challenge and the TUT Urban Acoustic Scenes 2018 dataset provided for the task, and evaluates the performance of a baseline system in the task.
-
19 Jun 2018 2 repositories listedIn this paper, we propose a system that consists of a simple fusion of two methods of the aforementioned types: a deep learning approach where log-scaled mel-spectrograms are input to a convolutional neural network, and…
-
3 May 2025 1 repository listedThis paper presents the Low-Complexity Acoustic Scene Classification with Device Information Task of the DCASE 2025 Challenge and its baseline system.
-
12 Jun 2024 1 repository listedThis work is an improved system that we submitted to task 1 of DCASE2023 challenge.
-
16 May 2024 1 repository listedThis article describes the Data-Efficient Low-Complexity Acoustic Scene Classification Task in the DCASE 2024 Challenge and the corresponding baseline system.
-
5 Feb 2024 1 repository listedIn addition, considering the abundance of unlabeled acoustic scene data in the real world, it is important to study the possible ways to utilize these unlabelled data.
-
2 Feb 2024 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Augmenting large language models (LLMs) to understand audio -- including non-speech sounds and non-verbal speech -- is critically important for diverse real-world applications of LLMs.
-
21 Nov 2023 1 repository listedThis paper presents AudioLog, a large language models (LLMs)-powered audio logging system with hybrid token-semantic contrastive learning.
-
5 Oct 2023 1 repository listedThe results show the feasibility of recognizing diverse acoustic scenes based on the audio event-relational graph.
-
28 Sep 2023 1 repository listedThe correlation between the sharpness of loss minima and generalisation in the context of deep neural networks has been subject to discussion for a long time.
-
13 Jun 2023 1 repository listedIn the Acoustic Scene Classification task (ASC), domain shift is mainly caused by different recording devices.
-
12 May 2023 1 repository listedHowever, we also show that DIR augmentation and Freq-MixStyle are complementary, achieving a new state-of-the-art performance on signals recorded by devices unseen during training.
-
3 May 2023 1 repository listedIn this paper, we study unsupervised approaches to improve the learning framework of such representations with unpaired text and audio.
-
4 Nov 2022 1 repository listedThis paper describes a pipeline for collecting acoustic scene data by using crowdsourcing.
-
27 Oct 2022 1 repository listedExperiments on a polyphonic acoustic scene dataset show that the proposed ERGL achieves competitive performance on ASC by using only a limited number of embeddings of audio events without any data augmentations.
-
27 Oct 2022 1 repository listedHowever, the computational complexity of computing the pairwise similarity matrix is high, particularly when a convolutional layer has many filters.
-
29 Mar 2022 1 repository listedWe propose a passive filter pruning framework, where a few convolutional filters from the CNNs are eliminated to yield compressed CNNs.
-
26 Dec 2021 1 repository listedThe approach used not only challenges some of the fundamental mathematical techniques used so far in early experiments of the same trend but also introduces new scopes and new horizons for interesting results.
-
26 Oct 2021 1 repository listedThe deployment of machine listening algorithms in real-life applications is often impeded by a domain shift caused for instance by different microphone characteristics.
-
16 Oct 2021 1 repository listedWe propose a variational Bayesian (VB) approach to learning distributions of latent variables in deep neural network (DNN) models for cross-domain knowledge transfer, to address acoustic mismatches between training and…
-
28 May 2021 1 repository listedThe most used techniques among the submissions were residual networks and weight quantization, with the top systems reaching over 70% accuracy, and log loss under 0.
-
26 May 2021 1 repository listedAs state-of-the-art CNN architectures-in computer vision and other domains-tend to go deeper in terms of number of layers, their RF size increases and therefore they degrade in performance in several audio…
-
25 May 2021 1 repository listedThis method works for both time and frequency domain representations of audio recordings.
-
5 Nov 2020 1 repository listedDeep Neural Networks are known to be very demanding in terms of computing and memory requirements.
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections