Methods › General › Self-Supervised Learning › M2D

Masked Modeling Duo

M2D

7 papers tagged archive 2025-07-28

Introduced by Daisuke Niizumi et al. in Masked Modeling Duo: Towards a Universal Audio Pre-training Framework

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

Self-supervised learning (SSL) using masked prediction has made great strides in general-purpose audio representation. This study proposes Masked Modeling Duo (M2D), an improved masked prediction SSL, which learns by predicting representations of masked input signals that serve as training signals. Unlike conventional methods, M2D obtains a training signal by encoding only the masked part, encouraging the two networks in M2D to model the input. While M2D improves general-purpose audio representations, a specialized representation is essential for real-world applications, such as in industrial and medical domains. The often confidential and proprietary data in such domains is typically limited in size and has a different distribution from that in pre-training datasets. Therefore, we propose M2D for X (M2D-X), which extends M2D to enable the pre-training of specialized representations for an application X. M2D-X learns from M2D and an additional task and inputs background noise. We make the additional task configurable to serve diverse applications, while the background noise helps learn on small data and forms a denoising task that makes representation robust. With these design choices, M2D-X should learn a representation specialized to serve various application needs. Our experiments confirmed that the representations for general-purpose audio, specialized for the highly competitive AudioSet and speech domain, and a small-data medical task achieve top-level performance, demonstrating the potential of using our models as a universal audio pre-training framework.

PaperSource

Papers archive 2025-07-28

7 shown of 7, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 28 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Self-Supervised Learning6
Audio Classification4
Audio Tagging2
Denoising2
Linear evaluation2
Music Genre Classification2
Speaker Identification2
Transfer Learning2
Audio captioning1
Audio to Text Retrieval1
Classify murmurs1
Emotion Recognition1
Environment Sound Classification1
GPU1
Instrument Recognition1
Keyword Spotting1
Keyword Spotting on Google Speech Commands1
Knowledge Distillation1
Music Auto-Tagging1
Music Classification1

Usage over time archive 2025-07-28

Papers per year tagged with M2D: 2022 to 2025, peak 4 4 0 2022: 1 paper 2022 2023: 1 paper 2023 2024: 4 papers 2024 2025: 1 paper 2025
Papers per year the archive tags with this method, by the paper's archive date (7 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Self-Supervised Learning

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections