Papers › AUTOMATED AUDIO CAPTIONING BY FINE-TUNING BART WITH AUDIOSET TAGS
AUTOMATED AUDIO CAPTIONING BY FINE-TUNING BART WITH AUDIOSET TAGS
F ́elix Gontier, Romain Serizel, Christophe Cerisara
utomated audio captioning is the multimodal task of describing environmental audio recordings with fluent natural language. Most current methods utilize pre-trained analysis models to extract rele- vant semantic content from the audio input. However, prior infor- mation on language modeling is rarely introduced, and correspond- ing architectures are limited in capacity due to data scarcity. In this paper, we present a method leveraging the linguistic informa- tion contained in BART, a large-scale conditional language model with general purpose pre-training. The caption generation is condi- tioned on sequences of textual AudioSet tags. This input is enriched with temporally aligned audio embeddings that allows the model to improve the sound event recognition. The full BART architecture is fine-tuned with few additional parameters. Experimental results demonstrate that, beyond the scaling properties of the architecture, language-only pre-training improves the text quality in the multi- modal setting of audio captioning. The best model achieves state- of-the-art performance on AudioCaps with 46.5 SPIDEr.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
1 archive task tag without a task page not shown.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Audio captioning | AudioCaps | BART + YAMNet + PANNs | CIDEr | 0.753 | #14 of 18 | Archive leaderboard | report |
| Audio captioning | AudioCaps | BART + YAMNet + PANNs | SPICE | 0.176 | #14 of 18 | Archive leaderboard | report |
| Audio captioning | AudioCaps | BART + YAMNet + PANNs | SPIDEr | 0.465 | #14 of 18 | Archive leaderboard | report |
| Retrieval-augmented Few-shot In-context Audio Captioning | AudioCaps | Automated audio captioning by fine-tuning bart with audioset tags | CIDEr | 0.147 | #5 of 5 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections