{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/noisy-activation-functions","title":"Noisy Activation Functions","arxiv_id":"1603.00391","date":"2016-03-01","proceeding":null,"authors":["Caglar Gulcehre","Marcin Moczulski","Misha Denil","Yoshua Bengio"],"abstract":"Common nonlinear activation functions used in neural networks can cause\ntraining difficulties due to the saturation behavior of the activation\nfunction, which may hide dependencies that are not visible to vanilla-SGD\n(using first order gradients only). Gating mechanisms that use softly\nsaturating activation functions to emulate the discrete switching of digital\nlogic circuits are good examples of this. We propose to exploit the injection\nof appropriate noise so that the gradients may flow easily, even if the\nnoiseless application of the activation function would yield zero gradient.\nLarge noise will dominate the noise-free gradient and allow stochastic gradient\ndescent toexplore more. By adding noise only to the problematic parts of the\nactivation function, we allow the optimization procedure to explore the\nboundary between the degenerate (saturating) and the well-behaved parts of the\nactivation function. We also establish connections to simulated annealing, when\nthe amount of noise is annealed down, making it easier to optimize hard\nobjective functions. We find experimentally that replacing such saturating\nactivation functions by noisy variants helps training in many contexts,\nyielding state-of-the-art or competitive results on different datasets and\ntask, especially when training seems to be the most difficult, e.g., when\ncurriculum learning is necessary to obtain good results.","url_abs":"http://arxiv.org/abs/1603.00391v3","url_pdf":"http://arxiv.org/pdf/1603.00391v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"noisy-activation-functions","repo_url":"https://github.com/wojciechz/learning_to_execute","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"torch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1603.00391","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}