Papers › Audiovisual Masked Autoencoders
Audiovisual Masked Autoencoders
Mariana-Iuliana Georgescu, Eduardo Fonseca, Radu Tudor Ionescu, Mario Lucic, Cordelia Schmid, Anurag Arnab
Can we leverage the audiovisual information already present in video to improve self-supervised representation learning? To answer this question, we study various pretraining architectures and objectives within the masked autoencoding framework, motivated by the success of similar methods in natural language and image understanding. We show that we can achieve significant improvements on audiovisual downstream classification tasks, surpassing the state-of-the-art on VGGSound and AudioSet. Furthermore, we can leverage our audiovisual pretraining scheme for multiple unimodal downstream tasks using a single audiovisual pretrained model. We additionally demonstrate the transferability of our representations, achieving state-of-the-art audiovisual results on Epic Kitchens without pretraining specifically for this dataset.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Audio Classification | AudioSet | Audiovisual Masked Autoencoder (Audiovisual, Single) | Test mAP | 0.518 | #5 of 51 | Archive leaderboard | report |
| Audio Classification | AudioSet | Audiovisual Masked Autoencoder (Audio-only, Single) | Test mAP | 0.466 | #35 of 51 | Archive leaderboard | report |
| Audio Classification | EPIC-KITCHENS-100 | Audiovisual Masked Autoencoder (Audiovisual, Single) | Top-1 Action | 46.0 | #1 of 4 | Archive leaderboard | report |
| Audio Classification | EPIC-KITCHENS-100 | Audiovisual Masked Autoencoder (Audiovisual, Single) | Top-1 Noun | 56.4 | #1 of 4 | Archive leaderboard | report |
| Audio Classification | EPIC-KITCHENS-100 | Audiovisual Masked Autoencoder (Audiovisual, Single) | Top-1 Verb | 71.4 | #1 of 4 | Archive leaderboard | report |
| Audio Classification | EPIC-KITCHENS-100 | Audiovisual Masked Autoencoder (Video-only, Single) | Top-1 Action | 45.8 | #2 of 4 | Archive leaderboard | report |
| Audio Classification | EPIC-KITCHENS-100 | Audiovisual Masked Autoencoder (Video-only, Single) | Top-1 Noun | 55.9 | #2 of 4 | Archive leaderboard | report |
| Audio Classification | EPIC-KITCHENS-100 | Audiovisual Masked Autoencoder (Video-only, Single) | Top-1 Verb | 70.8 | #2 of 4 | Archive leaderboard | report |
| Audio Classification | EPIC-KITCHENS-100 | Audiovisual Masked Autoencoder (Audio-only, Single) | Top-1 Action | 19.7 | #3 of 4 | Archive leaderboard | report |
| Audio Classification | EPIC-KITCHENS-100 | Audiovisual Masked Autoencoder (Audio-only, Single) | Top-1 Noun | 27.2 | #3 of 4 | Archive leaderboard | report |
| Audio Classification | EPIC-KITCHENS-100 | Audiovisual Masked Autoencoder (Audio-only, Single) | Top-1 Verb | 52.7 | #3 of 4 | Archive leaderboard | report |
| Audio Classification | VGGSound | Audiovisual Masked Autoencoder (Audiovisual, Single) | Top 1 Accuracy | 65.0 | #10 of 23 | Archive leaderboard | report |
| Audio Classification | VGGSound | Audiovisual Masked Autoencoder (Audio-only, Single) | Top 1 Accuracy | 57.2 | #14 of 23 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections