Datasets › AudioSet

AudioSet

Introduced in Audio Set: An ontology and human-labeled dataset for audio events1 Jan 2017 archive 2025-07-28

Audioset is an audio event dataset, which consists of over 2M human-annotated 10-second video clips. These clips are collected from YouTube, therefore many of which are in poor-quality and contain multiple sound-sources. A hierarchical ontology of 632 event classes is employed to annotate these data, which means that the same sound could be annotated as different labels. For example, the sound of barking is annotated as Animal, Pets, and Dog. All the videos are split into Evaluation/Balanced-Train/Unbalanced-Train set.

Source: Curriculum Audiovisual Learning

Benchmarks archive 2025-07-28

All 5 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

30 shown of 38 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 744. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
SSLAM: Enhancing Self-Supervised Models with Audio Mixtures for Polyphonic Soundscapes 1 1 13 Jun 2025 ran 5 of 7 samples (2 unverified)
M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP 2 1 28 Mar 2025 not harvested
DASS: Distilled Audio State Space Models Are Stronger and More Duration-Scalable Learners 1 2 4 Jul 2024 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation 2 1 4 Jun 2024 not harvested
MAX-AST: COMBINING CONVOLUTION, LOCAL AND GLOBAL SELF-ATTENTIONS FOR AUDIO EVENT CLASSIFICATION 1 1 14 Apr 2024 not harvested
Masked Modeling Duo: Towards a Universal Audio Pre-training Framework 2 2 9 Apr 2024 ran 7 of 7 samples (0 unverified; 7 pointer-only for licence)
DTF-AT: Decoupled Time-Frequency Audio Transformer for Event Classification 1 1 24 Mar 2024 not harvested
EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning 1 1 14 Mar 2024 ran 6 of 13 samples (7 unverified)
CLAPSep: Leveraging Contrastive Pre-trained Model for Multi-Modal Query-Conditioned Target Sound Extraction 1 1 27 Feb 2024 ran 4 of 4 samples (0 unverified)
EAT: Self-Supervised Pre-Training with Efficient Audio Transformer 1 1 7 Jan 2024 ran 13 of 17 samples (4 unverified)
OmniVec2 - A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning 0 1 1 Jan 2024 not harvested
OmniVec: Learning robust representations with cross modal sharing 0 1 7 Nov 2023 not harvested
Dynamic Convolutional Neural Networks as Efficient Pre-trained Audio Models 1 2 24 Oct 2023 not harvested
Self-supervised Audio Teacher-Student Transformer for Both Clip-level and Frame-level Tasks 2 2 7 Jun 2023 not harvested
BEATs: Audio Pre-Training with Acoustic Tokenizers 4 2 18 Dec 2022 ran 3 of 18 samples (15 unverified)
Audiovisual Masked Autoencoders 2 2 9 Dec 2022 not harvested
Efficient Large-scale Audio Tagging via Transformer-to-CNN Knowledge Distillation 2 4 9 Nov 2022 ran 1 of 2 samples (1 unverified)
Play It Back: Iterative Attention for Audio Recognition 1 1 20 Oct 2022 not harvested
Contrastive Audio-Visual Masked Autoencoder 1 6 2 Oct 2022 ran 0 of 1 samples (1 unverified)
UAVM: Towards Unifying Audio and Visual Models 1 2 29 Jul 2022 ran 6 of 14 samples (8 unverified)
End-to-End Audio Strikes Back: Boosting Augmentations Towards An Efficient Audio Classification Network 1 2 25 Apr 2022 not harvested
HTS-AT: A Hierarchical Token-Semantic Audio Transformer for Sound Classification and Detection 1 1 2 Feb 2022 ran 4 of 9 samples (5 unverified)
Zero-shot Audio Source Separation through Query-based Learning from Weakly-labeled Data 1 2 15 Dec 2021 ran 3 of 10 samples (7 unverified)
Conformer-Based Self-Supervised Learning for Non-Speech Audio Tasks 0 1 14 Oct 2021 not harvested
Efficient Training of Audio Transformers with Patchout 2 3 11 Oct 2021 ran 3 of 3 samples (0 unverified)
Attention Bottlenecks for Multimodal Fusion 1 1 30 Jun 2021 not harvested
VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text 5 1 22 Apr 2021 ran 5 of 8 samples (3 unverified; 8 pointer-only for licence)
AST: Audio Spectrogram Transformer 5 3 5 Apr 2021 ran 0 of 3 samples (3 unverified)
Multi-Format Contrastive Learning of Audio Representations 0 1 11 Mar 2021 not harvested
Perceiver: General Perception with Iterative Attention 12 1 4 Mar 2021 ran 41 of 55 samples (14 unverified; 11 pointer-only for licence)

The full list of 38 is in the JSON twin.

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • AudioSet

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections