Datasets › AudioCaps

AudioCaps

Introduced by Chris Dongjoo Kim et al. in AudioCaps: Generating Captions for Audios in The Wild1 Jun 2019 archive 2025-07-28

AudioCaps is a dataset of sounds with event descriptions that was introduced for the task of audio captioning, with sounds sourced from the AudioSet dataset. Annotators were provided the audio tracks together with category hints (and with additional video hints if needed).

Source: Audio Retrieval with Natural Language Queries

Image source: https://audiocaps.github.io/

Benchmarks archive 2025-07-28

All 6 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

30 shown of 43 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 279. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport 2 1 16 Jan 2025 not harvested
TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization 1 2 30 Dec 2024 not harvested
ETTA: Elucidating the Design Space of Text-to-Audio Models 1 2 26 Dec 2024 not harvested
Enhancing Retrieval-Augmented Audio Captioning with Generation-Assisted Multimodal Querying and Progressive Learning 0 1 14 Oct 2024 not harvested
SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs 1 1 12 Oct 2024 not harvested
EnCLAP++: Analyzing the EnCLAP Framework for Optimizing Automated Audio Captioning Performance 1 2 2 Sep 2024 not harvested
Stable Audio Open 1 1 19 Jul 2024 ran 1 of 2 samples (1 unverified)
Taming Data and Transformers for Audio Generation 1 2 27 Jun 2024 not harvested
Enhancing Automated Audio Captioning via Large Language Models with Optimized Audio Encoding 1 1 19 Jun 2024 ran 5 of 8 samples (3 unverified)
Improving Text-To-Audio Models with Synthetic Captions 1 1 18 Jun 2024 ran 3 of 4 samples (1 unverified; 4 pointer-only for licence)
Long-form music generation with latent diffusion 1 1 16 Apr 2024 not harvested
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding 2 1 22 Mar 2024 not harvested
CLAPSep: Leveraging Contrastive Pre-trained Model for Multi-Modal Query-Conditioned Target Sound Extraction 1 1 27 Feb 2024 ran 4 of 4 samples (0 unverified)
Fast Timing-Conditioned Latent Audio Diffusion 2 1 7 Feb 2024 ran 2 of 4 samples (2 unverified)
Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities 1 2 2 Feb 2024 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
EnCLAP: Combining Neural Audio Codec and Audio-Text Joint Embedding for Automated Audio Captioning 1 2 31 Jan 2024 not harvested
Auffusion: Leveraging the Power of Diffusion and Large Language Models for Text-to-Audio Generation 1 2 2 Jan 2024 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
Audiobox: Unified Audio Generation with Natural Language Prompts 0 1 25 Dec 2023 not harvested
Zero-shot audio captioning with audio-language model guidance and audio context keywords 1 2 14 Nov 2023 ran 13 of 17 samples (4 unverified; 17 pointer-only for licence)
ConsistencyTTA: Accelerating Diffusion-Based Text-to-Audio Generation with Consistency Distillation 1 1 19 Sep 2023 ran 4 of 4 samples (0 unverified)
RECAP: Retrieval-Augmented Audio Captioning 1 1 18 Sep 2023 ran 5 of 8 samples (3 unverified; 8 pointer-only for licence)
Retrieval-Augmented Text-to-Audio Generation 0 1 14 Sep 2023 not harvested
Zero-Shot Audio Captioning via Audibility Guidance 0 1 7 Sep 2023 not harvested
Rethinking Transfer and Auxiliary Learning for Improving Audio Captioning Transformer 0 1 20 Aug 2023 not harvested
AudioLDM 2: Learning Holistic Audio Generation with Self-supervised Pretraining 2 2 10 Aug 2023 ran 12 of 27 samples (15 unverified; 13 pointer-only for licence)
VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset 2 2 29 May 2023 ran 15 of 42 samples (27 unverified)
Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation 1 1 29 May 2023 ran 8 of 12 samples (4 unverified)
Any-to-Any Generation via Composable Diffusion 2 1 19 May 2023 ran 2 of 13 samples (11 unverified)
ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities 2 1 18 May 2023 ran 2 of 7 samples (5 unverified)
Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model 1 1 24 Apr 2023 not harvested

The full list of 43 is in the JSON twin.

Dataset loaders archive 2025-07-28

3 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • AudioCaps

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections