Methods › General › Model Compression › Soups
Model Soups
Soups
Introduced by Mitchell Wortsman et al. in Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Compress an ensemble of models into a single one by averaging their weights (under certain pre-conditions).
Papers archive 2025-07-28
19 shown of 19, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
UC-MOA: Utility-Conditioned Multi-Objective Alignment for Distributional Pareto-Optimality 10 Mar 2025 · 0 repositories · arXiv:2503.10669
-
Local Superior Soups: A Catalyst for Model Merging in Cross-Silo Federated Learning 31 Oct 2024 · 1 repository · arXiv:2410.23660Syntology ran 3 of 12 samples · 9 unverified · 12 pointer-only (licence)
-
Merging in a Bottle: Differentiable Adaptive Merging (DAM) and the Path from Averaging to Automation 10 Oct 2024 · 1 repository · arXiv:2410.08371
-
Robust Biharmonic Skinning Using Geometric Fields 1 Jun 2024 · 0 repositories · arXiv:2406.00238
-
Cross-Dataset Generalization For Retinal Lesions Segmentation 14 May 2024 · 0 repositories · arXiv:2405.08329
-
How Much You Ate? Food Portion Estimation on Spoons 12 May 2024 · 0 repositories · arXiv:2405.08717
-
FissionFusion: Fast Geometric Generation and Hierarchical Souping for Medical Image Analysis 20 Mar 2024 · 1 repository · arXiv:2403.13341
-
RADIN: Souping on a Budget 31 Jan 2024 · 0 repositories · arXiv:2401.17790
-
Partial Fine-Tuning: A Successor to Full Fine-Tuning for Vision Transformers 25 Dec 2023 · 0 repositories · arXiv:2312.15681
-
Descriptor and Word Soups: Overcoming the Parameter Efficiency Accuracy Tradeoff for Out-of-Distribution Few-shot Learning 21 Nov 2023 · 1 repository · arXiv:2311.13612
-
Sparse Model Soups: A Recipe for Improved Pruning via Model Averaging 29 Jun 2023 · 1 repository · arXiv:2306.16788
-
Graph Ladling: Shockingly Simple Parallel GNN Training without Intermediate Communication 18 Jun 2023 · 1 repository · arXiv:2306.10466Syntology ran 1 of 1 samples · 0 unverified
-
Pre-training Language Model as a Multi-perspective Course Learner 6 May 2023 · 0 repositories · arXiv:2305.03981
-
Seasoning Model Soups for Robustness to Adversarial and Natural Distribution Shifts 20 Feb 2023 · 0 repositories · arXiv:2302.10164
-
Candidate Soups: Fusing Candidate Results Improves Translation Quality for Non-Autoregressive Translation 27 Jan 2023 · 1 repository · arXiv:2301.11503
-
Model soups to increase inference without increasing compute time 24 Jan 2023 · 1 repository · arXiv:2301.10092
-
Revisiting adapters with adversarial training 10 Oct 2022 · 0 repositories · arXiv:2210.04886
-
Bag of Tricks for Domain Adaptive Multi-Object Tracking 31 May 2022 · 1 repository · arXiv:2205.15609
-
Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time 10 Mar 2022 · 6 repositories · arXiv:2203.05482Syntology ran 5 of 17 samples · 12 unverified
Tasks archive 2025-07-28
20 shown of 22 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections