Browse State-of-the-Art › Audio Tagging
Audio Tagging
49 papers with code · 1 benchmark · 10 datasets archive 2025-07-28
Audio tagging is a task to predict the tags of audio clips. Audio tagging tasks include music tagging, acoustic scene classification, audio event classification, etc.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| AudioSet (11 rows) | CAV-MAE (Audio-Visual) | Contrastive Audio-Visual Masked Autoencoder | code | Syntology ran 0 of 1 samples · 1 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
10 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 49 papers with code (81 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
5 Apr 2021 5 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedIn the past decade, convolutional neural networks (CNNs) have been widely adopted as the main building block for end-to-end audio classification models, which aim to learn a direct mapping from audio spectrograms to…
-
27 Jun 2018 5 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedWe present an end-to-end deep learning approach to denoising speech signals by processing the raw waveform directly.
-
14 Sep 2019 4 repositories listed Syntology ran 1 of 8 samples · 7 unverifiedPronounced as "musician", the musicnn library contains a set of pre-trained musically motivated convolutional neural networks for music audio tagging: https://github.
-
26 Jul 2018 3 repositories listedThe goal of the task is to build an audio tagging system that can recognize the category of an audio clip from a subset of 41 diverse categories drawn from the AudioSet Ontology.
-
28 Mar 2025 2 repositories listedIn the second stage, it learns CLAP features using the audio features learned from the LLM-based embeddings.
-
25 Sep 2023 2 repositories listedDilated convolution with learnable spacings (DCLS) is a recent convolution method in which the positions of the kernel elements are learned throughout training by backpropagation.
-
7 Jun 2023 2 repositories listedIn order to tackle both clip-level and frame-level tasks, this paper proposes Audio Teacher-Student Transformer (ATST), with a clip-level version (named ATST-Clip) and a frame-level version (named ATST-Frame),…
-
9 Nov 2022 2 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedWe provide models of different complexity levels, scaling from low-complexity models up to a new state-of-the-art performance of .
-
11 Oct 2021 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedHowever, one of the main shortcomings of transformer models, compared to the well-established CNNs, is the computational complexity.
-
24 Aug 2020 2 repositories listedThis paper describes CRNNs we used to participate in Task 5 of the DCASE 2020 challenge.
-
7 Jun 2019 2 repositories listedThe task evaluates systems for multi-label audio tagging using a large set of noisy-labeled data, and a much smaller set of manually-labeled data, under a large vocabulary setting of 80 everyday sound classes.
-
30 Oct 2018 2 repositories listedAudio tagging is challenging due to the limited size of data and noisy labels.
-
27 Jun 2018 2 repositories listedWe present an end-to-end deep learning approach to denoising speech signals by processing the raw waveform directly.
-
24 Feb 2017 2 repositories listedIn this paper, we propose to use a convolutional neural network (CNN) to extract robust features from mel-filter banks (MFBs), spectrograms or even raw waveforms for audio tagging.
-
13 Jul 2016 2 repositories listedFor the unsupervised feature learning, we propose to use a symmetric or asymmetric deep de-noising auto-encoder (sDAE or aDAE) to generate new data-driven features from the Mel-Filter Banks (MFBs) features.
-
28 Mar 2025 1 repository listedImmersive communication has made significant advancements, especially with the release of the codec for Immersive Voice and Audio Services.
-
26 Mar 2025 1 repository listedAudioSet is one of the most used and largest datasets in audio tagging, containing about 2 million audio samples that are manually labeled with 527 event categories organized into an ontology.
-
19 Mar 2025 1 repository listedLarge Language Models (LLMs) have recently shown remarkable ability to process not only text but also multimodal inputs such as speech and audio.
-
14 Mar 2025 1 repository listedWe target the problem of developing new low-complexity networks for the sound event detection task.
-
17 Feb 2025 1 repository listedTherefore, we propose a new method: MAsked latenT Prediction And Classification (MATPAC), which is trained with two pretext tasks solved jointly.
-
17 Sep 2024 1 repository listedWe introduce Label-Combination Prototypical Networks (LC-Protonets) to address the problem of multi-label few-shot classification, where a model must generalize to new classes based on only a few available examples.
-
22 Jul 2024 1 repository listedThe broadcasting industry is increasingly adopting IP techniques, revolutionising both live and pre-recorded content production, from news gathering to live music events.
-
4 Jul 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)To the best of our knowledge, it is the first SSM that outperforms the Transformers on AudioSet and achieves an mAP of 48.
-
18 Dec 2023 1 repository listedIn the age of music streaming platforms, the task of automatically tagging music audio has garnered significant attention, driving researchers to devise methods aimed at enhancing performance metrics on standard…
-
24 Oct 2023 1 repository listedAudio Spectrogram Transformers are excellent at exploiting large datasets, creating powerful pre-trained models that surpass CNNs when fine-tuned on downstream tasks.
-
15 Jun 2023 1 repository listedIn this paper, we analyze how the performance of large-scale pretrained audio neural networks designed for audio pattern recognition changes when deployed on a hardware such as Raspberry Pi.
-
30 May 2023 1 repository listedSounds carry an abundance of information about activities and events in our everyday environment, such as traffic noise, road works, music, or people talking.
-
16 Apr 2023 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)However, such semantic consistency from the synchronization is hard to guarantee in unconstrained videos, due to the irrelevant modality noise and differentiated semantic correlation.
-
23 Jan 2023 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 1 pointer-only (licence)Attention-based models are appealing for multimodal processing because inputs from multiple modalities can be concatenated and fed to a single backbone network - thus requiring very little fusion engineering.
-
22 Nov 2022 1 repository listedThe proposed metric, ontology-aware mean average precision (OmAP) addresses the weaknesses of mAP by utilizing the AudioSet ontology information during the evaluation.
Syntology lines on 8 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections