{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/group-sparse-regularization-for-deep-neural","title":"Group Sparse Regularization for Deep Neural Networks","arxiv_id":"1607.00485","date":"2016-07-02","proceeding":null,"authors":["Simone Scardapane","Danilo Comminiello","Amir Hussain","Aurelio Uncini"],"abstract":"In this paper, we consider the joint task of simultaneously optimizing (i)\nthe weights of a deep neural network, (ii) the number of neurons for each\nhidden layer, and (iii) the subset of active input features (i.e., feature\nselection). While these problems are generally dealt with separately, we\npresent a simple regularized formulation allowing to solve all three of them in\nparallel, using standard optimization routines. Specifically, we extend the\ngroup Lasso penalty (originated in the linear regression literature) in order\nto impose group-level sparsity on the network's connections, where each group\nis defined as the set of outgoing weights from a unit. Depending on the\nspecific case, the weights can be related to an input variable, to a hidden\nneuron, or to a bias unit, thus performing simultaneously all the\naforementioned tasks in order to obtain a compact network. We perform an\nextensive experimental evaluation, by comparing with classical weight decay and\nLasso penalties. We show that a sparse version of the group Lasso penalty is\nable to achieve competitive performances, while at the same time resulting in\nextremely compact networks with a smaller number of input features. We evaluate\nboth on a toy dataset for handwritten digit recognition, and on multiple\nrealistic large-scale classification problems.","url_abs":"http://arxiv.org/abs/1607.00485v1","url_pdf":"http://arxiv.org/pdf/1607.00485v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"group-sparse-regularization-for-deep-neural","repo_url":"https://bitbucket.org/ispamm/group-lasso-deep-networks","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"handwritten-digit-recognition","task_name":"Handwritten Digit Recognition"},{"task_slug":"feature-selection","task_name":"feature selection"}],"methods":[{"method_slug":"linear-regression","method_name":"Linear Regression"},{"method_slug":"weight-decay","method_name":"Weight Decay"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1607.00485","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}