{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/evolutionary-stochastic-gradient-descent-for","title":"Evolutionary Stochastic Gradient Descent for Optimization of Deep Neural Networks","arxiv_id":"1810.06773","date":"2018-10-16","proceeding":"NeurIPS 2018 12","authors":["Xiaodong Cui","Wei zhang","Zoltán Tüske","Michael Picheny"],"abstract":"We propose a population-based Evolutionary Stochastic Gradient Descent (ESGD)\nframework for optimizing deep neural networks. ESGD combines SGD and\ngradient-free evolutionary algorithms as complementary algorithms in one\nframework in which the optimization alternates between the SGD step and\nevolution step to improve the average fitness of the population. With a\nback-off strategy in the SGD step and an elitist strategy in the evolution\nstep, it guarantees that the best fitness in the population will never degrade.\nIn addition, individuals in the population optimized with various SGD-based\noptimizers using distinct hyper-parameters in the SGD step are considered as\ncompeting species in a coevolution setting such that the complementarity of the\noptimizers is also taken into account. The effectiveness of ESGD is\ndemonstrated across multiple applications including speech recognition, image\nrecognition and language modeling, using networks with a variety of deep\narchitectures.","url_abs":"http://arxiv.org/abs/1810.06773v1","url_pdf":"http://arxiv.org/pdf/1810.06773v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"evolutionary-stochastic-gradient-descent-for","repo_url":"https://github.com/tqch/esgd-ws","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"evolutionary-algorithms","task_name":"Evolutionary Algorithms"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[{"method_slug":"sgd","method_name":"SGD"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1810.06773","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}