Papers › M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP

M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP

28 Mar 2025arXiv:2503.22104archive 2025-07-28

Daisuke Niizumi, Daiki Takeuchi, Masahiro Yasuda, Binh Thien Nguyen, Yasunori Ohishi, Noboru Harada

Contrastive language-audio pre-training (CLAP) has addressed audio-language tasks such as audio-text retrieval by aligning audio and text in a common feature space. While CLAP addresses general audio-language tasks, its audio features do not generalize well in audio tasks. In contrast, self-supervised learning (SSL) models learn general-purpose audio features that perform well in diverse audio tasks. We pursue representation learning that can be widely used in audio applications and hypothesize that a method that learns both general audio features and CLAP features should achieve our goal, which we call a general-purpose audio-language representation. To implement our hypothesis, we propose M2D2, a second-generation masked modeling duo (M2D) that combines an SSL M2D and CLAP. M2D2 learns two types of features using two modalities (audio and text) in a two-stage training process. It also utilizes advanced LLM-based sentence embeddings in CLAP training for powerful semantic supervision. In the first stage, M2D2 learns generalizable audio features from M2D and CLAP, where CLAP aligns the features with the fine LLM-based semantic embeddings. In the second stage, it learns CLAP features using the audio features learned from the LLM-based embeddings. Through these pre-training stages, M2D2 should enhance generalizability and performance in its audio and CLAP features. Experiments validated that M2D2 achieves effective general-purpose audio-language representation, highlighted with SOTA fine-tuning mAP of 49.0 for AudioSet, SOTA performance in music tasks, and top-level performance in audio-language tasks.

PaperPDFCode

Code

nttcslab/eval-audio-repr pytorchNOASSERTION report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Audio ClassificationAudio TaggingAudio captioningAudio to Text RetrievalEmotion RecognitionInstrument RecognitionMusic Auto-TaggingMusic ClassificationMusic Genre ClassificationMusic TaggingRepresentation LearningSelf-Supervised LearningSentence EmbeddingsSinger IdentificationText RetrievalText to Audio RetrievalVocal technique classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Audio Classification AudioSet M2D2 Test mAP 0.490 #15 of 51 Archive leaderboard report
Audio Classification ESC-50 M2D2 AS+ Accuracy (5-fold) 98.5 #3 of 29 Archive leaderboard report
Audio Classification ESC-50 M2D2 AS+ PRE-TRAINING DATASET AudioSet,WavCaps #3 of 29 Archive leaderboard report
Audio Classification ESC-50 M2D2 AS+ Top-1 Accuracy 98.5 #3 of 29 Archive leaderboard report
Emotion Recognition Emomusic M2D-CLAP EmoA 77.4 #1 of 5 Archive leaderboard report
Emotion Recognition Emomusic M2D-CLAP EmoV 61.9 #1 of 5 Archive leaderboard report
Emotion Recognition Emomusic M2D2 EmoA 76.7 #2 of 5 Archive leaderboard report
Emotion Recognition Emomusic M2D2 EmoV 59.3 #2 of 5 Archive leaderboard report
Emotion Recognition Emomusic M2D EmoA 76.1 #3 of 5 Archive leaderboard report
Emotion Recognition Emomusic M2D EmoV 59.4 #3 of 5 Archive leaderboard report
Instrument Recognition NSynth M2D-CLAP Accuracy 80.6 #1 of 7 Archive leaderboard report
Instrument Recognition NSynth M2D2 AS+ Accuracy 79.7 #2 of 7 Archive leaderboard report
Instrument Recognition NSynth M2D AS Accuracy 78.7 #3 of 7 Archive leaderboard report
Music Auto-Tagging MagnaTagATune M2D2 AS+ PR-AUC 41.6 #1 of 3 Archive leaderboard report
Music Auto-Tagging MagnaTagATune M2D2 AS+ ROC AUC 91.8 #1 of 3 Archive leaderboard report
Singer Identification VocalSet M2D2 AS+ Accuracy 92.7 #1 of 2 Archive leaderboard report
Singer Identification VocalSet M2D2 Accuracy 91.8 #2 of 2 Archive leaderboard report
Vocal technique classification VocalSet M2D2 AS+ Accuracy 78.9 #1 of 2 Archive leaderboard report
Vocal technique classification VocalSet M2D2 Accuracy 77.4 #2 of 2 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

M2D

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections