Papers › LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport

16 Jan 2025arXiv:2501.09291archive 2025-07-28

Kyeongha Rho, Hyeongkeun Lee, Valentio Iverson, Joon Son Chung

Automated audio captioning is a task that generates textual descriptions for audio content, and recent studies have explored using visual information to enhance captioning quality. However, current methods often fail to effectively fuse audio and visual data, missing important semantic cues from each modality. To address this, we introduce LAVCap, a large language model (LLM)-based audio-visual captioning framework that effectively integrates visual information with audio to improve audio captioning performance. LAVCap employs an optimal transport-based alignment loss to bridge the modality gap between audio and visual features, enabling more effective semantic extraction. Additionally, we propose an optimal transport attention module that enhances audio-visual fusion using an optimal transport assignment map. Combined with the optimal training strategy, experimental results demonstrate that each component of our framework is effective. LAVCap outperforms existing state-of-the-art methods on the AudioCaps dataset, without relying on large datasets or post-processing. Code is available at https://github.com/NAVER-INTEL-Co-Lab/gaudi-lavcap.

PaperPDFCode

Code

naver-intel-co-lab/gaudi-lavcap officialmentioned in papermentioned on GitHubpytorch report
Hyeongkeun/LAVCap mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Audio captioningImage CaptioningLanguage ModelingLanguage ModellingLarge Language Model

2 archive task tags without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Audio captioning AudioCaps LAVCap BLEU-4 0.297 #3 of 18 Archive leaderboard report
Audio captioning AudioCaps LAVCap CIDEr 0.849 #3 of 18 Archive leaderboard report
Audio captioning AudioCaps LAVCap METEOR 0.262 #3 of 18 Archive leaderboard report
Audio captioning AudioCaps LAVCap ROUGE-L 0.510 #3 of 18 Archive leaderboard report
Audio captioning AudioCaps LAVCap SPICE 0.185 #3 of 18 Archive leaderboard report
Audio captioning AudioCaps LAVCap SPIDEr 0.517 #3 of 18 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AttentionSoftmax

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections