Papers › MT-SLVR: Multi-Task Self-Supervised Learning for Transformation In(Variant) Representations
MT-SLVR: Multi-Task Self-Supervised Learning for Transformation In(Variant) Representations
Calum Heggan, Tim Hospedales, Sam Budgett, Mehrdad Yaghoobi
Contrastive self-supervised learning has gained attention for its ability to create high-quality representations from large unlabelled data sets. A key reason that these powerful features enable data-efficient learning of downstream tasks is that they provide augmentation invariance, which is often a useful inductive bias. However, the amount and type of invariances preferred is not known apriori, and varies across different downstream tasks. We therefore propose a multi-task self-supervised framework (MT-SLVR) that learns both variant and invariant features in a parameter-efficient manner. Our multi-task representation provides a strong and flexible feature that benefits diverse downstream tasks. We evaluate our approach on few-shot classification tasks drawn from a variety of audio domains and demonstrate improved classification performance on all of them
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Few-Shot Audio Classification | BirdClef 2020 (Pruned) | SimCLR (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 30.93±0.38 | #8 of 10 | Archive leaderboard | report |
| Few-Shot Audio Classification | BirdClef 2020 (Pruned) | MT-SLVR (SimCLR + MLAP) w/ Parallel Adapters (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 29.49±0.38 | #9 of 10 | Archive leaderboard | report |
| Few-Shot Audio Classification | BirdClef 2020 (Pruned) | Multi-Label Augmentation Prediction (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 21.04±0.35 | #10 of 10 | Archive leaderboard | report |
| Few-Shot Audio Classification | CREMA-D | MT-SLVR (SimCLR + MLAP) w/ Parallel Adapters (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 29.61±0.38 | #1 of 3 | Archive leaderboard | report |
| Few-Shot Audio Classification | CREMA-D | SimCLR (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 29.10±0.36 | #2 of 3 | Archive leaderboard | report |
| Few-Shot Audio Classification | CREMA-D | Multi-Label Augmentation Prediction (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 21.68±0.33 | #3 of 3 | Archive leaderboard | report |
| Few-Shot Audio Classification | Common Voice | MT-SLVR (SimCLR + MLAP) w/ Parallel Adapters (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 35.22±0.40 | #1 of 3 | Archive leaderboard | report |
| Few-Shot Audio Classification | Common Voice | SimCLR (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 33.33±0.38 | #2 of 3 | Archive leaderboard | report |
| Few-Shot Audio Classification | Common Voice | Multi-Label Augmentation Prediction (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 23.00±0.42 | #3 of 3 | Archive leaderboard | report |
| Few-Shot Audio Classification | ESC-50 | MT-SLVR (SimCLR + MLAP) w/ Parallel Adapters (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 69.53±0.39 | #4 of 10 | Archive leaderboard | report |
| Few-Shot Audio Classification | ESC-50 | SimCLR (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 63.40±0.39 | #8 of 10 | Archive leaderboard | report |
| Few-Shot Audio Classification | ESC-50 | Multi-Label Augmentation Prediction (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 37.76±0.34 | #10 of 10 | Archive leaderboard | report |
| Few-Shot Audio Classification | FSDKaggle2018 | MT-SLVR (SimCLR + MLAP) w/ Parallel Adapters (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 39.11±0.41 | #6 of 10 | Archive leaderboard | report |
| Few-Shot Audio Classification | FSDKaggle2018 | SimCLR (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 37.64±0.40 | #8 of 10 | Archive leaderboard | report |
| Few-Shot Audio Classification | FSDKaggle2018 | Multi-Label Augmentation Prediction (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 21.72±0.34 | #10 of 10 | Archive leaderboard | report |
| Few-Shot Audio Classification | NSynth | MT-SLVR (SimCLR + MLAP) w/ Parallel Adapters (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 71.81±0.39 | #6 of 10 | Archive leaderboard | report |
| Few-Shot Audio Classification | NSynth | SimCLR (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 66.44±0.40 | #8 of 10 | Archive leaderboard | report |
| Few-Shot Audio Classification | NSynth | Multi-Label Augmentation Prediction (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 62.52±0.36 | #10 of 10 | Archive leaderboard | report |
| Few-Shot Audio Classification | Speech Accent Archive | MT-SLVR (SimCLR + MLAP) w/ Parallel Adapters (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 28.92±0.37 | #1 of 3 | Archive leaderboard | report |
| Few-Shot Audio Classification | Speech Accent Archive | SimCLR (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 26.16±0.34 | #2 of 3 | Archive leaderboard | report |
| Few-Shot Audio Classification | Speech Accent Archive | Multi-Label Augmentation Prediction (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 23.08±0.34 | #3 of 3 | Archive leaderboard | report |
| Few-Shot Audio Classification | Speech Command v2 | SimCLR (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 25.68±0.35 | #1 of 3 | Archive leaderboard | report |
| Few-Shot Audio Classification | Speech Command v2 | MT-SLVR (SimCLR + MLAP) w/ Parallel Adapters (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 23.65±0.34 | #2 of 3 | Archive leaderboard | report |
| Few-Shot Audio Classification | Speech Command v2 | Multi-Label Augmentation Prediction (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 20.08±0.37 | #3 of 3 | Archive leaderboard | report |
| Few-Shot Audio Classification | VoxCeleb1 | MT-SLVR (SimCLR + MLAP) w/ Parallel Adapters (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 33.58±0.39 | #6 of 10 | Archive leaderboard | report |
| Few-Shot Audio Classification | VoxCeleb1 | SimCLR (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 31.18±0.37 | #7 of 10 | Archive leaderboard | report |
| Few-Shot Audio Classification | VoxCeleb1 | Multi-Label Augmentation Prediction (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 21.68±0.40 | #10 of 10 | Archive leaderboard | report |
| Few-Shot Audio Classification | Watkins Marine Mammal Sounds | MT-SLVR (SimCLR + MLAP) w/ Parallel Adapters (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 59.49±0.42 | #1 of 5 | Archive leaderboard | report |
| Few-Shot Audio Classification | Watkins Marine Mammal Sounds | SimCLR (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 52.91±0.41 | #3 of 5 | Archive leaderboard | report |
| Few-Shot Audio Classification | Watkins Marine Mammal Sounds | Multi-Label Augmentation Prediction (FSD50K, RN18) | Top-1 Accuracy(5-Way-1-Shot) | 28.88±0.39 | #5 of 5 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections