Papers › Masked Latent Prediction and Classification for Self-Supervised Audio Representation Learning

Masked Latent Prediction and Classification for Self-Supervised Audio Representation Learning

17 Feb 2025ICASSP 2025 3arXiv:2502.12031archive 2025-07-28

Aurian Quelennec, Pierre Chouteau, Geoffroy Peeters, Slim Essid

Recently, self-supervised learning methods based on masked latent prediction have proven to encode input data into powerful representations. However, during training, the learned latent space can be further transformed to extract higher-level information that could be more suited for downstream classification tasks. Therefore, we propose a new method: MAsked latenT Prediction And Classification (MATPAC), which is trained with two pretext tasks solved jointly. As in previous work, the first pretext task is a masked latent prediction task, ensuring a robust input representation in the latent space. The second one is unsupervised classification, which utilises the latent representations of the first pretext task to match probability distributions between a teacher and a student. We validate the MATPAC method by comparing it to other state-of-the-art proposals and conducting ablations studies. MATPAC reaches state-of-the-art self-supervised learning results on reference audio classification datasets such as OpenMIC, GTZAN, ESC-50 and US8K and outperforms comparable supervised methods results for musical auto-tagging on Magna-tag-a-tune.

PaperPDFConference PDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Audio ClassificationAudio TaggingClassificationEnvironmental Sound ClassificationInstrument RecognitionMusic Auto-TaggingMusic Genre ClassificationMusic TaggingPredictionRepresentation LearningSelf-Supervised Audio ClassificationSelf-Supervised LearningTAG

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Audio Classification ESC-50 MATPAC (SSL model, linear eval) Accuracy (5-fold) 93.5 #18 of 29 Archive leaderboard report
Audio Classification ESC-50 MATPAC (SSL model, linear eval) PRE-TRAINING DATASET AudioSet #18 of 29 Archive leaderboard report
Audio Classification ESC-50 MATPAC (SSL model, linear eval) Top-1 Accuracy 93.5 #18 of 29 Archive leaderboard report
Audio Classification FSD50K MATPAC (SSL Model) mAP 55.2 #7 of 10 Archive leaderboard report
Environmental Sound Classification UrbanSound8K MATPAC (SSL, linear eval) Accuracy 89.4 #2 of 3 Archive leaderboard report
Instrument Recognition NSynth MATPAC (SSL, linear eval) Accuracy 74.6 #4 of 7 Archive leaderboard report
Instrument Recognition OpenMIC-2018 MATPAC (SSL Model, linear eval) mean average precision 0.854 #2 of 5 Archive leaderboard report
Music Auto-Tagging MagnaTagATune MATPAC (SSL, linear eval) PR-AUC 41.1 #2 of 3 Archive leaderboard report
Music Auto-Tagging MagnaTagATune MATPAC (SSL, linear eval) ROC AUC 91.6 #2 of 3 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections