Papers › CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition
CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition
Jongseo Lee, Joohyun Chang, DongHo Lee, Jinwoo Choi
We propose Cross-Attention in Audio, Space, and Time (CA^2ST), a transformer-based method for holistic video recognition. Recognizing actions in videos requires both spatial and temporal understanding, yet most existing models lack a balanced spatio-temporal understanding of videos. To address this, we propose a novel two-stream architecture, called Cross-Attention in Space and Time (CAST), using only RGB input. In each layer of CAST, Bottleneck Cross-Attention (B-CA) enables spatial and temporal experts to exchange information and make synergistic predictions. For holistic video understanding, we extend CAST by integrating an audio expert, forming Cross-Attention in Visual and Audio (CAVA). We validate the CAST on benchmarks with different characteristics, EPIC-KITCHENS-100, Something-Something-V2, and Kinetics-400, consistently showing balanced performance. We also validate the CAVA on audio-visual action recognition benchmarks, including UCF-101, VGG-Sound, KineticsSound, and EPIC-SOUNDS. With a favorable performance of CAVA across these datasets, we demonstrate the effective information exchange among multiple experts within the B-CA module. In summary, CA^2ST combines CAST and CAVA by employing spatial, temporal, and audio experts through cross-attention, achieving balanced and holistic video understanding.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Action Classification | Kinetics-Sounds | CA2ST(B/16) | Top 1 Accuracy | 93.3 | #1 of 4 | Archive leaderboard | report |
| Action Classification | Kinetics-Sounds | CAVA(B/16) | Top 1 Accuracy | 92.9 | #2 of 4 | Archive leaderboard | report |
| Action Recognition | UCF101 | CA2ST(B/16) | 3-fold Accuracy | 97.2 | #22 of 91 | Archive leaderboard | report |
| Audio Classification | EPIC-SOUNDS | CA2ST(B/16) | Accuracy | 61 | #2 of 3 | Archive leaderboard | report |
| Audio Classification | EPIC-SOUNDS | CAVA(B/16) | Accuracy | 60.3 | #3 of 3 | Archive leaderboard | report |
| Audio Classification | VGGSound | CA2ST(B/16) | Top 1 Accuracy | 68.3 | #2 of 23 | Archive leaderboard | report |
| Audio Classification | VGGSound | CAVA(B/16) | Top 1 Accuracy | 68.2 | #4 of 23 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections