Datasets › VGG-Sound

VGG-Sound

Introduced by Honglie Chen et al. in VGGSound: A Large-scale Audio-Visual Dataset archive 2025-07-28

Consists of more than 210k videos for 310 audio classes.

Source: VGGSound: A Large-scale Audio-Visual Dataset

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Audio Classification VGGSound Mirasol3B Top 1 Accuracy 69.8 Mirasol3B: A Multimodal Autoregressive model for... — 23 Compare
Video-to-Sound Generation VGG-Sound MMAudio-S-16kHz FAD 0.79 MMAudio: Taming Multimodal Joint Training for... hkchengrex/MMAudio 8 Compare
Multi-modal Classification VGG-Sound MMT Top-1 Accuracy 66.2 Multiscale Multimodal Transformer for Multimodal Action... — 4 Compare

Papers archive 2025-07-28

19 shown of 19 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 211. The Syntology column is from Syntology's graph (read 2026-09-25), stated per sample; it is not part of any archive number.

DateSamples run Syntology
CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition 0 2 30 Mar 2025 not harvested
MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis 1 2 19 Dec 2024 ran 4 of 4 samples (0 unverified)
Tell What You Hear From What You See -- Video to Audio Generation Through Text 1 1 8 Nov 2024 ran 0 of 6 samples (6 unverified; 6 pointer-only for licence)
Temporally Aligned Audio for Video with Autoregression 1 1 20 Sep 2024 ran 11 of 13 samples (2 unverified)
Masked Generative Video-to-Audio Transformers with Enhanced Synchronicity 0 1 15 Jul 2024 not harvested
Read, Watch and Scream! Sound Generation from Text and Video 1 1 8 Jul 2024 ran 18 of 22 samples (4 unverified; 22 pointer-only for licence)
Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching 1 1 1 Jun 2024 ran 14 of 17 samples (3 unverified)
EquiAV: Leveraging Equivariance for Audio-Visual Contrastive Learning 1 1 14 Mar 2024 ran 6 of 13 samples (7 unverified)
Mirasol3B: A Multimodal Autoregressive model for time-aligned and contextual modalities 0 1 9 Nov 2023 not harvested
V2A-Mapper: A Lightweight Solution for Vision-to-Audio Generation by Connecting Foundation Models 1 1 18 Aug 2023 not harvested
ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities 2 2 18 May 2023 ran 2 of 7 samples (5 unverified)
Multiscale Audio Spectrogram Transformer for Efficient Audio Classification 0 1 19 Mar 2023 not harvested
Audiovisual Masked Autoencoders 2 2 9 Dec 2022 not harvested
Play It Back: Iterative Attention for Audio Recognition 1 1 20 Oct 2022 not harvested
Contrastive Audio-Visual Masked Autoencoder 1 3 2 Oct 2022 ran 0 of 1 samples (1 unverified)
Multiscale Multimodal Transformer for Multimodal Action Recognition 0 3 22 Sep 2022 not harvested
AVT: Audio-Video Transformer for Multimodal Action Recognition 0 3 22 Sep 2022 not harvested
UAVM: Towards Unifying Audio and Visual Models 1 4 29 Jul 2022 ran 10 of 14 samples (4 unverified)
Attention Bottlenecks for Multimodal Fusion 1 3 30 Jun 2021 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

CC-BY 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • VGG-Sound
  • VGGSound

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections