Papers › Attention Bottlenecks for Multimodal Fusion

Attention Bottlenecks for Multimodal Fusion

30 Jun 2021NeurIPS 2021 12arXiv:2107.00135archive 2025-07-28

Arsha Nagrani, Shan Yang, Anurag Arnab, Aren Jansen, Cordelia Schmid, Chen Sun

Humans perceive the world by concurrently processing and fusing high-dimensional inputs from multiple modalities such as vision and audio. Machine perception models, in stark contrast, are typically modality-specific and optimised for unimodal benchmarks, and hence late-stage fusion of final representations or predictions from each modality (`late-fusion') is still a dominant paradigm for multimodal video classification. Instead, we introduce a novel transformer based architecture that uses `fusion bottlenecks' for modality fusion at multiple layers. Compared to traditional pairwise self-attention, our model forces information between different modalities to pass through a small number of bottleneck latents, requiring the model to collate and condense the most relevant information in each modality and only share what is necessary. We find that such a strategy improves fusion performance, at the same time reducing computational cost. We conduct thorough ablation studies, and achieve state-of-the-art results on multiple audio-visual classification benchmarks including Audioset, Epic-Kitchens and VGGSound. All code and models will be released.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action ClassificationAction RecognitionAudio ClassificationVideo Classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Classification Kinetics-400 MBT (AV) Acc@1 80.8 #92 of 207 Archive leaderboard report
Action Classification Kinetics-400 MBT (AV) Acc@5 94.6 #92 of 207 Archive leaderboard report
Action Classification Kinetics-Sounds MBT (AV) Top 1 Accuracy 85 #4 of 4 Archive leaderboard report
Action Classification Kinetics-Sounds MBT (AV) Top 5 Accuracy 96.8 #4 of 4 Archive leaderboard report
Action Classification MiT MBT (AV) Top 1 Accuracy 37.3 #13 of 29 Archive leaderboard report
Action Classification MiT MBT (AV) Top 5 Accuracy 61.2 #13 of 29 Archive leaderboard report
Action Recognition EPIC-KITCHENS-100 MBT Action@1 43.4 #24 of 32 Archive leaderboard report
Action Recognition EPIC-KITCHENS-100 MBT Noun@1 58 #24 of 32 Archive leaderboard report
Action Recognition EPIC-KITCHENS-100 MBT Verb@1 64.8 #24 of 32 Archive leaderboard report
Audio Classification AudioSet MBT (AS-500K training + Video) Test mAP 0.496 #12 of 51 Archive leaderboard report
Audio Classification VGGSound MBT (A) Top 1 Accuracy 52.3 #20 of 23 Archive leaderboard report
Audio Classification VGGSound MBT (A) Top 5 Accuracy 78.1 #20 of 23 Archive leaderboard report
Audio Classification VGGSound MBT (V) Top 1 Accuracy 51.2 #21 of 23 Archive leaderboard report
Audio Classification VGGSound MBT (V) Top 5 Accuracy 72.6 #21 of 23 Archive leaderboard report
Audio Classification VGGSound MBT (AV) Top 5 Accuracy 85.6 #23 of 23 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections