Papers › Audiovisual Masked Autoencoders

Audiovisual Masked Autoencoders

9 Dec 2022ICCV 2023 1arXiv:2212.05922archive 2025-07-28

Mariana-Iuliana Georgescu, Eduardo Fonseca, Radu Tudor Ionescu, Mario Lucic, Cordelia Schmid, Anurag Arnab

Can we leverage the audiovisual information already present in video to improve self-supervised representation learning? To answer this question, we study various pretraining architectures and objectives within the masked autoencoding framework, motivated by the success of similar methods in natural language and image understanding. We show that we can achieve significant improvements on audiovisual downstream classification tasks, surpassing the state-of-the-art on VGGSound and AudioSet. Furthermore, we can leverage our audiovisual pretraining scheme for multiple unimodal downstream tasks using a single audiovisual pretrained model. We additionally demonstrate the transferability of our representations, achieving state-of-the-art audiovisual results on Epic Kitchens without pretraining specifically for this dataset.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

google-research/scenic officialmentioned in papermentioned on GitHubjax report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Audio ClassificationRepresentation Learning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Audio Classification AudioSet Audiovisual Masked Autoencoder (Audiovisual, Single) Test mAP 0.518 #5 of 51 Archive leaderboard report
Audio Classification AudioSet Audiovisual Masked Autoencoder (Audio-only, Single) Test mAP 0.466 #35 of 51 Archive leaderboard report
Audio Classification EPIC-KITCHENS-100 Audiovisual Masked Autoencoder (Audiovisual, Single) Top-1 Action 46.0 #1 of 4 Archive leaderboard report
Audio Classification EPIC-KITCHENS-100 Audiovisual Masked Autoencoder (Audiovisual, Single) Top-1 Noun 56.4 #1 of 4 Archive leaderboard report
Audio Classification EPIC-KITCHENS-100 Audiovisual Masked Autoencoder (Audiovisual, Single) Top-1 Verb 71.4 #1 of 4 Archive leaderboard report
Audio Classification EPIC-KITCHENS-100 Audiovisual Masked Autoencoder (Video-only, Single) Top-1 Action 45.8 #2 of 4 Archive leaderboard report
Audio Classification EPIC-KITCHENS-100 Audiovisual Masked Autoencoder (Video-only, Single) Top-1 Noun 55.9 #2 of 4 Archive leaderboard report
Audio Classification EPIC-KITCHENS-100 Audiovisual Masked Autoencoder (Video-only, Single) Top-1 Verb 70.8 #2 of 4 Archive leaderboard report
Audio Classification EPIC-KITCHENS-100 Audiovisual Masked Autoencoder (Audio-only, Single) Top-1 Action 19.7 #3 of 4 Archive leaderboard report
Audio Classification EPIC-KITCHENS-100 Audiovisual Masked Autoencoder (Audio-only, Single) Top-1 Noun 27.2 #3 of 4 Archive leaderboard report
Audio Classification EPIC-KITCHENS-100 Audiovisual Masked Autoencoder (Audio-only, Single) Top-1 Verb 52.7 #3 of 4 Archive leaderboard report
Audio Classification VGGSound Audiovisual Masked Autoencoder (Audiovisual, Single) Top 1 Accuracy 65.0 #10 of 23 Archive leaderboard report
Audio Classification VGGSound Audiovisual Masked Autoencoder (Audio-only, Single) Top 1 Accuracy 57.2 #14 of 23 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections