Papers › AVT: Audio-Video Transformer for Multimodal Action Recognition

AVT: Audio-Video Transformer for Multimodal Action Recognition

22 Sep 2022Submitted to ICLR 2022 9archive 2025-07-28

Wentao Zhu, Jingru Yi, Kevin Hsu, Xiaohang Sun, Xiang Hao, Linda Liu, Mohamed Omar

Action recognition is an essential field for video understanding. To learn from heterogeneous data sources effectively, in this work, we propose a novel multimodal action recognition approach termed Audio-Video Transformer (AVT). AVT uses a combination of video and audio signals to improve action recognition accuracy, leveraging the effective spatio-temporal representation by the video Transformer. For multimodal fusion, simply concatenating multimodal tokens in a cross-modal Transformer requires large computational and memory resources, instead we reduce the cross-modality complexity through an audio-video bottleneck Transformer. To improve the learning efficiency of multimodal Transformer, we integrate self-supervised objectives, i.e., audio-video contrastive learning, audio-video matching, and masked audio and video learning, into AVT training, which maps diverse audio and video representations into a common multimodal representation space. We further propose a masked audio segment loss to learn semantic audio activities in AVT. Extensive experiments and ablation studies on three public datasets and two in-house datasets consistently demonstrate the effectiveness of the proposed AVT. Specifically, AVT outperforms its previous state-of-the-art counterparts on Kinetics-Sounds and Epic-Kitchens-100 datasets by 8% and 1%, respectively, without external training data. AVT also surpasses one of the previous state-of-the-art video Transformers by 10% on the VGGSound dataset by leveraging the audio signal. Compared to one of the previous state-of-the-art multimodal Transformers, AVT is 1.3x more efficient in terms of FLOPs and improves the accuracy by 4.2% on Epic-Kitchens-100. Visualization results further demonstrate that the audio provides complementary and discriminative features, and our AVT can effectively understand the action from a combination of audio and video.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionAudio ClassificationContrastive LearningMulti-modal ClassificationVideo Understanding

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Recognition EPIC-KITCHENS-100 AVT Action@1 47.2 #15 of 32 Archive leaderboard report
Action Recognition EPIC-KITCHENS-100 AVT Noun@1 59.3 #15 of 32 Archive leaderboard report
Action Recognition EPIC-KITCHENS-100 AVT Verb@1 70.4 #15 of 32 Archive leaderboard report
Audio Classification VGGSound AVT (Audio-Visual) Top 1 Accuracy 63.9 #11 of 23 Archive leaderboard report
Audio Classification VGGSound AVT (Audio-Visual) Top 5 Accuracy 85.0 #11 of 23 Archive leaderboard report
Audio Classification VGGSound AVT (V) Top 1 Accuracy 53.2 #19 of 23 Archive leaderboard report
Audio Classification VGGSound AVT (V) Top 5 Accuracy 74.8 #19 of 23 Archive leaderboard report
Multi-modal Classification VGG-Sound AVT Top-1 Accuracy 63.9 #4 of 4 Archive leaderboard report
Multi-modal Classification VGG-Sound AVT Top-5 Accuracy 85.0 #4 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections