Methods › General › Self-Supervised Learning › M2D
Masked Modeling Duo
M2D
Introduced by Daisuke Niizumi et al. in Masked Modeling Duo: Towards a Universal Audio Pre-training Framework
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Self-supervised learning (SSL) using masked prediction has made great strides in general-purpose audio representation. This study proposes Masked Modeling Duo (M2D), an improved masked prediction SSL, which learns by predicting representations of masked input signals that serve as training signals. Unlike conventional methods, M2D obtains a training signal by encoding only the masked part, encouraging the two networks in M2D to model the input. While M2D improves general-purpose audio representations, a specialized representation is essential for real-world applications, such as in industrial and medical domains. The often confidential and proprietary data in such domains is typically limited in size and has a different distribution from that in pre-training datasets. Therefore, we propose M2D for X (M2D-X), which extends M2D to enable the pre-training of specialized representations for an application X. M2D-X learns from M2D and an additional task and inputs background noise. We make the additional task configurable to serve diverse applications, while the background noise helps learn on small data and forms a denoising task that makes representation robust. With these design choices, M2D-X should learn a representation specialized to serve various application needs. Our experiments confirmed that the representations for general-purpose audio, specialized for the highly competitive AudioSet and speech domain, and a small-data medical task achieve top-level performance, demonstrating the potential of using our models as a universal audio pre-training framework.
Papers archive 2025-07-28
7 shown of 7, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
M2D2: Exploring General-purpose Audio-Language Representations Beyond CLAP 28 Mar 2025 · 2 repositories · arXiv:2503.22104
-
BOXR: Body and head motion Optimization framework for eXtended Reality 16 Oct 2024 · 1 repository · arXiv:2410.13084
-
M2D-CLAP: Masked Modeling Duo Meets CLAP for Learning General-purpose Audio-Language Representation 4 Jun 2024 · 2 repositories · arXiv:2406.02032
-
Exploring Pre-trained General-purpose Audio Representations for Heart Murmur Detection 26 Apr 2024 · 2 repositories · arXiv:2404.17107
-
Masked Modeling Duo: Towards a Universal Audio Pre-training Framework 9 Apr 2024 · 2 repositories · arXiv:2404.06095Syntology ran 7 of 7 samples · 0 unverified · 7 pointer-only (licence)
-
Masked Modeling Duo for Speech: Specializing General-Purpose Audio Representation to Speech using Denoising Distillation 23 May 2023 · 1 repository · arXiv:2305.14079
-
Masked Modeling Duo: Learning Representations by Encouraging Both Networks to Model the Input 26 Oct 2022 · 1 repository · arXiv:2210.14648Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)
Tasks archive 2025-07-28
20 shown of 28 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
| Task | Papers |
|---|---|
| Self-Supervised Learning | 6 |
| Audio Classification | 4 |
| Audio Tagging | 2 |
| Denoising | 2 |
| Linear evaluation | 2 |
| Music Genre Classification | 2 |
| Speaker Identification | 2 |
| Transfer Learning | 2 |
| Audio captioning | 1 |
| Audio to Text Retrieval | 1 |
| Classify murmurs | 1 |
| Emotion Recognition | 1 |
| Environment Sound Classification | 1 |
| GPU | 1 |
| Instrument Recognition | 1 |
| Keyword Spotting | 1 |
| Keyword Spotting on Google Speech Commands | 1 |
| Knowledge Distillation | 1 |
| Music Auto-Tagging | 1 |
| Music Classification | 1 |
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections