Papers › AUTOMATED AUDIO CAPTIONING BY FINE-TUNING BART WITH AUDIOSET TAGS

AUTOMATED AUDIO CAPTIONING BY FINE-TUNING BART WITH AUDIOSET TAGS

15 Nov 2021DCASE workshop 2021 11archive 2025-07-28

F ́elix Gontier, Romain Serizel, Christophe Cerisara

utomated audio captioning is the multimodal task of describing environmental audio recordings with fluent natural language. Most current methods utilize pre-trained analysis models to extract rele- vant semantic content from the audio input. However, prior infor- mation on language modeling is rarely introduced, and correspond- ing architectures are limited in capacity due to data scarcity. In this paper, we present a method leveraging the linguistic informa- tion contained in BART, a large-scale conditional language model with general purpose pre-training. The caption generation is condi- tioned on sequences of textual AudioSet tags. This input is enriched with temporally aligned audio embeddings that allows the model to improve the sound event recognition. The full BART architecture is fine-tuned with few additional parameters. Experimental results demonstrate that, beyond the scaling properties of the architecture, language-only pre-training improves the text quality in the multi- modal setting of audio captioning. The best model achieves state- of-the-art performance on AudioCaps with 46.5 SPIDEr.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Audio captioningCaption GenerationLanguage ModelingLanguage ModellingRetrieval-augmented Few-shot In-context Audio Captioning

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Audio captioning AudioCaps BART + YAMNet + PANNs CIDEr 0.753 #14 of 18 Archive leaderboard report
Audio captioning AudioCaps BART + YAMNet + PANNs SPICE 0.176 #14 of 18 Archive leaderboard report
Audio captioning AudioCaps BART + YAMNet + PANNs SPIDEr 0.465 #14 of 18 Archive leaderboard report
Retrieval-augmented Few-shot In-context Audio Captioning AudioCaps Automated audio captioning by fine-tuning bart with audioset tags CIDEr 0.147 #5 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AdamAttentionBARTBPEDense ConnectionsDropoutLayer NormalizationLinear LayerMulti-Head AttentionResidual ConnectionSoftmax

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections