{"url":"/method/snapshot-ensembles","slug":"snapshot-ensembles","name":"Snapshot Ensembles","full_name":"Snapshot Ensembles: Train 1, get M for free","full_name_withheld":false,"description_markdown":"The  overhead  cost  of  training  multiple  deep  neural networks  could  be  very  high  in  terms  of  the  training  time, hardware, and computational resource requirement and often acts  as  obstacle  for  creating  deep  ensembles.  To  overcome these barriers Huang et al. proposed a unique method to create  ensemble  which  at  the  cost  of  training  one  model, yields  multiple  constituent  model  snapshots  that  can  be ensembled together to create a strong learner. The core idea behind the concept is to make the model converge to several local minima along the optimization path and save the model parameters at these local minima points. During the training phase, a neural network would traverse through many such points. The lowest of all such local minima is known as the Global Minima. The larger the model, more are the number of parameters and larger the number of local minima points. This implies, there are discrete sets of weights and biases, at which  the  model  is  making  fewer  errors.  So,  every  such minimum  can  be  considered a  weak  but  a  potential learner model for the problem being solved. Multiple such snapshot of  weights  and  biases  are  recorded  which  can  later  be ensembled to get a better generalized model which makes the least amount of mistakes.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Snapshot Ensembles: Train 1, get M for free","paper":"/paper/snapshot-ensembles-train-1-get-m-for-free","first_author":"Gao Huang","n_authors":6,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/snapshot-ensembles-train-1-get-m-for-free"},"source":{"url":"http://arxiv.org/abs/1704.00109v1","title":"Snapshot Ensembles: Train 1, get M for free","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Active Learning","url":"/methods/category/active-learning","pwc_aliases":[]}],"n_papers_tagged":8,"archive_num_papers":8,"papers_newest_first":[{"paper":null,"title":"SnapE -- Training Snapshot Ensembles of Link Prediction Models","date":"2024-08-05","arxiv_id":"2408.02707","n_code_links":0,"syntology":null},{"paper":"/paper/to-stay-or-not-to-stay-in-the-pre-train-basin","title":"To Stay or Not to Stay in the Pre-train Basin: Insights on Ensembling in Transfer Learning","date":"2023-03-06","arxiv_id":"2303.03374","n_code_links":1,"syntology":{"ran":2,"of":2,"unverified":0,"pointer_only":2}},{"paper":null,"title":"Interpretable Diversity Analysis: Visualizing Feature Representations In Low-Cost Ensembles","date":"2023-02-12","arxiv_id":"2302.05822","n_code_links":0,"syntology":null},{"paper":"/paper/malaria-parasite-detection-using-efficient","title":"Malaria Parasite Detection using Efficient Neural Ensembles","date":"2021-10-15","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/accuracy-privacy-trade-off-in-deep-ensemble","title":"Accuracy-Privacy Trade-off in Deep Ensemble: A Membership Inference Perspective","date":"2021-05-12","arxiv_id":"2105.05381","n_code_links":1,"syntology":null},{"paper":null,"title":"MotherNets: Rapid Deep Ensemble Learning","date":"2018-09-12","arxiv_id":"1809.04270","n_code_links":0,"syntology":null},{"paper":"/paper/loss-surfaces-mode-connectivity-and-fast","title":"Loss Surfaces, Mode Connectivity, and Fast Ensembling of DNNs","date":"2018-02-27","arxiv_id":"1802.10026","n_code_links":8,"syntology":{"ran":3,"of":14,"unverified":11,"pointer_only":0}},{"paper":"/paper/snapshot-ensembles-train-1-get-m-for-free","title":"Snapshot Ensembles: Train 1, get M for free","date":"2017-04-01","arxiv_id":"1704.00109","n_code_links":11,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":0}}],"papers_shown":8,"tasks":[{"task":"/task/diversity","name":"Diversity","papers":2},{"task":"/task/ensemble-learning","name":"Ensemble Learning","papers":2},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":2},{"task":"/task/clustering","name":"Clustering","papers":1},{"task":"/task/clustering-ensemble","name":"Clustering Ensemble","papers":1},{"task":"/task/diagnostic","name":"Diagnostic","papers":1},{"task":"/task/inference-attack","name":"Inference Attack","papers":1},{"task":"/task/knowledge-distillation","name":"Knowledge Distillation","papers":1},{"task":"/task/knowledge-graphs","name":"Knowledge Graphs","papers":1},{"task":"/task/link-prediction","name":"Link Prediction","papers":1},{"task":"/task/medical-image-classification","name":"Medical Image Classification","papers":1},{"task":"/task/membership-inference-attack","name":"Membership Inference Attack","papers":1},{"task":"/task/prediction","name":"Prediction","papers":1}],"tasks_shown":13,"n_tasks":13,"usage_by_year":[{"year":"2017","papers":1},{"year":"2018","papers":2},{"year":"2021","papers":2},{"year":"2023","papers":2},{"year":"2024","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/snapshot-ensembles"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}