Browse State-of-the-Art › Audio Classification
Audio Classification
183 papers with code · 22 benchmarks · 43 datasets archive 2025-07-28
Audio Classification is a machine learning task that involves identifying and tagging audio signals into different classes or categories. The goal of audio classification is to enable machines to automatically recognize and distinguish between different types of audio, such as music, speech, and environmental sounds.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
25 leaderboard tables shown for this task, 22 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 25 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
43 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 43 until expanded.
Subtasks archive 2025-07-28
4 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 183 papers with code (361 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
29 Sep 2016 16 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedConvolutional Neural Networks (CNNs) have proven very effective in image classification and show promise for audio.
-
4 Mar 2021 12 repositories listed Syntology ran 41 of 55 samples · 14 unverified · 11 pointer-only (licence)The perception models used in deep learning on the other hand are designed for individual modalities, often relying on domain-specific assumptions such as the local grid structures exploited by virtually all existing…
-
3 Oct 2023 6 repositories listed Syntology ran 7 of 14 samples · 7 unverifiedWe thus propose VIDAL-10M with Video, Infrared, Depth, Audio and their corresponding Language, naming as VIDAL-10M.
-
22 Apr 2021 5 repositories listed Syntology ran 5 of 8 samples · 3 unverified · 8 pointer-only (licence)We train VATT end-to-end from scratch using multimodal contrastive losses and evaluate its performance by the downstream tasks of video action recognition, audio event classification, image classification, and…
-
5 Apr 2021 5 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedIn the past decade, convolutional neural networks (CNNs) have been widely adopted as the main building block for end-to-end audio classification models, which aim to learn a direct mapping from audio spectrograms to…
-
6 Mar 2018 5 repositories listed Syntology ran 0 of 11 samples · 11 unverifiedThe objective of audio classification is to predict the presence or absence of audio events in an audio clip.
-
18 Dec 2022 4 repositories listed Syntology ran 3 of 18 samples · 15 unverifiedIn the first iteration, we use random projection as the acoustic tokenizer to train an audio SSL model in a mask and label prediction manner.
-
13 Jul 2022 4 repositories listed Syntology ran 17 of 32 samples · 15 unverified · 10 pointer-only (licence)Following the Transformer encoder-decoder design in MAE, our Audio-MAE first encodes audio spectrogram patches with a high masking ratio, feeding only the non-masked tokens through encoder layers.
-
26 Apr 2022 4 repositories listedSelf-supervised learning (SSL) learns knowledge from a large amount of unlabeled data, and then transfers the knowledge to a specific problem with a limited number of labeled data.
-
21 Jan 2021 4 repositories listed Syntology ran 14 of 26 samples · 12 unverified · 2 pointer-only (licence)In this work we show that we can train a single learnable frontend that outperforms mel-filterbanks on a wide range of audio signals, including speech, music, audio events and animal sounds, providing a general-purpose…
-
15 Jun 2020 4 repositories listed Syntology ran 4 of 4 samples · 0 unverifiedBecause of the scale invariance, this modification only alters the effective step sizes without changing the effective update directions, thus enjoying the original convergence properties of GD optimizers.
-
25 Apr 2024 3 repositories listedMachine learning has the potential to revolutionize passive acoustic monitoring (PAM) for ecological assessments.
-
19 Oct 2021 3 repositories listed Syntology ran 11 of 16 samples · 5 unverifiedHowever, pure Transformer models tend to require more training data compared to CNNs, and the success of the AST relies on supervised pretraining that requires a large amount of labeled data and a complex training…
-
22 Jul 2020 3 repositories listedBesides, we show that even though we use the pretrained model weights for initialization, there is variance in performance in various output runs of the same model.
-
28 May 2019 3 repositories listed Syntology ran 0 of 13 samples · 13 unverifiedModern deep neural networks are well known to be brittle in the face of unknown data instances and recognition of the latter remains a challenge.
-
9 Jul 2018 3 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedExplainable Artificial Intelligence (XAI) is targeted at understanding how models perform feature selection and derive their classification decisions.
-
18 Feb 2016 3 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedTraditional convolutional layers extract features from patches of data by applying a non-linearity on an affine function of the input.
-
28 Mar 2025 2 repositories listedIn the second stage, it learns CLAP features using the audio features learned from the LLM-based embeddings.
-
4 Jun 2024 2 repositories listedHere, we explore a new representation, a general-purpose audio-language representation, that performs well in both ZS and transfer learning.
-
9 Apr 2024 2 repositories listed Syntology ran 7 of 7 samples · 0 unverified · 7 pointer-only (licence)This study proposes Masked Modeling Duo (M2D), an improved masked prediction SSL, which learns by predicting representations of masked input signals that serve as training signals.
-
22 Mar 2024 2 repositories listedWe introduce InternVideo2, a new family of video foundation models (ViFM) that achieve the state-of-the-art results in video recognition, video-text tasks, and video-centric dialogue.
-
30 Nov 2023 2 repositories listedIn this work, we introduce Acoustic Prompt Tuning (APT), a new adapter extending LLMs and VLMs to the audio domain by injecting audio embeddings to the input of LLMs, namely soft prompting.
-
14 Nov 2023 2 repositories listed Syntology ran 5 of 7 samples · 2 unverified · 7 pointer-only (licence)Recently, instruction-following audio-language models have received broad attention for audio interaction with humans.
-
25 Sep 2023 2 repositories listedDilated convolution with learnable spacings (DCLS) is a recent convolution method in which the positions of the kernel elements are learned throughout training by backpropagation.
-
12 Jul 2023 2 repositories listed Syntology ran 4 of 7 samples · 3 unverifiedWith the advent of deep learning models, classification of important signals from these datasets has markedly improved.
-
7 Jun 2023 2 repositories listedIn order to tackle both clip-level and frame-level tasks, this paper proposes Audio Teacher-Student Transformer (ATST), with a clip-level version (named ATST-Clip) and a frame-level version (named ATST-Frame),…
-
31 May 2023 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Compared to Perceiver IO, our model requires absolutely no modality-specific processing at inference time, and uses an order of magnitude fewer parameters at equivalent accuracy on ImageNet.
-
18 May 2023 2 repositories listed Syntology ran 2 of 7 samples · 5 unverifiedIn this work, we explore a scalable way for building a general representation model toward unlimited modalities.
-
9 Dec 2022 2 repositories listedCan we leverage the audiovisual information already present in video to improve self-supervised representation learning?
-
9 Nov 2022 2 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedWe provide models of different complexity levels, scaling from low-complexity models up to a new state-of-the-art performance of .
Syntology lines on 20 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections