{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/improving-deep-neural-network-with-multiple","title":"Improving Deep Neural Network with Multiple Parametric Exponential Linear Units","arxiv_id":"1606.00305","date":"2016-06-01","proceeding":null,"authors":["Yang Li","Chunxiao Fan","Yong Li","Qiong Wu","Yue Ming"],"abstract":"Activation function is crucial to the recent successes of deep neural\nnetworks. In this paper, we first propose a new activation function, Multiple\nParametric Exponential Linear Units (MPELU), aiming to generalize and unify the\nrectified and exponential linear units. As the generalized form, MPELU shares\nthe advantages of Parametric Rectified Linear Unit (PReLU) and Exponential\nLinear Unit (ELU), leading to better classification performance and convergence\nproperty. In addition, weight initialization is very important to train very\ndeep networks. The existing methods laid a solid foundation for networks using\nrectified linear units but not for exponential linear units. This paper\ncomplements the current theory and extends it to the wider range. Specifically,\nwe put forward a way of initialization, enabling training of very deep networks\nusing exponential linear units. Experiments demonstrate that the proposed\ninitialization not only helps the training process but leads to better\ngeneralization performance. Finally, utilizing the proposed activation function\nand initialization, we present a deep MPELU residual architecture that achieves\nstate-of-the-art performance on the CIFAR-10/100 datasets. The code is\navailable at https://github.com/Coldmooon/Code-for-MPELU.","url_abs":"http://arxiv.org/abs/1606.00305v3","url_pdf":"http://arxiv.org/pdf/1606.00305v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"improving-deep-neural-network-with-multiple","repo_url":"https://github.com/Coldmooon/Code-for-MPELU","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"torch","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}