{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/smoothout-smoothing-out-sharp-minima-to","title":"SmoothOut: Smoothing Out Sharp Minima to Improve Generalization in Deep Learning","arxiv_id":"1805.07898","date":"2018-05-21","proceeding":null,"authors":["Wei Wen","Yandan Wang","Feng Yan","Cong Xu","Chunpeng Wu","Yiran Chen","Hai Li"],"abstract":"In Deep Learning, Stochastic Gradient Descent (SGD) is usually selected as a\ntraining method because of its efficiency; however, recently, a problem in SGD\ngains research interest: sharp minima in Deep Neural Networks (DNNs) have poor\ngeneralization; especially, large-batch SGD tends to converge to sharp minima.\nIt becomes an open question whether escaping sharp minima can improve the\ngeneralization. To answer this question, we propose SmoothOut framework to\nsmooth out sharp minima in DNNs and thereby improve generalization. In a\nnutshell, SmoothOut perturbs multiple copies of the DNN by noise injection and\naverages these copies. Injecting noises to SGD is widely used in the\nliterature, but SmoothOut differs in lots of ways: (1) a de-noising process is\napplied before parameter updating; (2) noise strength is adapted to filter\nnorm; (3) an alternative interpretation on the advantage of noise injection,\nfrom the perspective of sharpness and generalization; (4) usage of uniform\nnoise instead of Gaussian noise. We prove that SmoothOut can eliminate sharp\nminima. Training multiple DNN copies is inefficient, we further propose an\nunbiased stochastic SmoothOut which only introduces the overhead of noise\ninjecting and de-noising per batch. An adaptive variant of SmoothOut,\nAdaSmoothOut, is also proposed to improve generalization. In a variety of\nexperiments, SmoothOut and AdaSmoothOut consistently improve generalization in\nboth small-batch and large-batch training on the top of state-of-the-art\nsolutions.","url_abs":"http://arxiv.org/abs/1805.07898v3","url_pdf":"http://arxiv.org/pdf/1805.07898v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"smoothout-smoothing-out-sharp-minima-to","repo_url":"https://github.com/wenwei202/smoothout","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"deep-learning","task_name":"Deep Learning"},{"task_slug":"open-question","task_name":"Open-Ended Question Answering"}],"methods":[{"method_slug":"sgd","method_name":"SGD"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.07898","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1805.07898"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/wenwei202/smoothout","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":1},"by_repo_kind":{"official":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"c02e8307733c548e","entry":"get_dataset","repo":"wenwei202/smoothout","repo_kind":"official","path":"data.py","file_url":"https://github.com/wenwei202/smoothout/blob/HEAD/data.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c02e8307733c548e"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}