Papers › CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition

CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition

30 Mar 2025arXiv:2503.23447archive 2025-07-28

Jongseo Lee, Joohyun Chang, DongHo Lee, Jinwoo Choi

We propose Cross-Attention in Audio, Space, and Time (CA^2ST), a transformer-based method for holistic video recognition. Recognizing actions in videos requires both spatial and temporal understanding, yet most existing models lack a balanced spatio-temporal understanding of videos. To address this, we propose a novel two-stream architecture, called Cross-Attention in Space and Time (CAST), using only RGB input. In each layer of CAST, Bottleneck Cross-Attention (B-CA) enables spatial and temporal experts to exchange information and make synergistic predictions. For holistic video understanding, we extend CAST by integrating an audio expert, forming Cross-Attention in Visual and Audio (CAVA). We validate the CAST on benchmarks with different characteristics, EPIC-KITCHENS-100, Something-Something-V2, and Kinetics-400, consistently showing balanced performance. We also validate the CAVA on audio-visual action recognition benchmarks, including UCF-101, VGG-Sound, KineticsSound, and EPIC-SOUNDS. With a favorable performance of CAVA across these datasets, we demonstrate the effective information exchange among multiple experts within the B-CA module. In summary, CA^2ST combines CAST and CAVA by employing spatial, temporal, and audio experts through cross-attention, achieving balanced and holistic video understanding.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action ClassificationAction RecognitionAudio ClassificationVideo RecognitionVideo Understanding

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Classification Kinetics-Sounds CA2ST(B/16) Top 1 Accuracy 93.3 #1 of 4 Archive leaderboard report
Action Classification Kinetics-Sounds CAVA(B/16) Top 1 Accuracy 92.9 #2 of 4 Archive leaderboard report
Action Recognition UCF101 CA2ST(B/16) 3-fold Accuracy 97.2 #22 of 91 Archive leaderboard report
Audio Classification EPIC-SOUNDS CA2ST(B/16) Accuracy 61 #2 of 3 Archive leaderboard report
Audio Classification EPIC-SOUNDS CAVA(B/16) Accuracy 60.3 #3 of 3 Archive leaderboard report
Audio Classification VGGSound CA2ST(B/16) Top 1 Accuracy 68.3 #2 of 23 Archive leaderboard report
Audio Classification VGGSound CAVA(B/16) Top 1 Accuracy 68.2 #4 of 23 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections