{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-state-of-sparsity-in-deep-neural-networks","title":"The State of Sparsity in Deep Neural Networks","arxiv_id":"1902.09574","date":"2019-02-25","proceeding":null,"authors":["Trevor Gale","Erich Elsen","Sara Hooker"],"abstract":"We rigorously evaluate three state-of-the-art techniques for inducing\nsparsity in deep neural networks on two large-scale learning tasks: Transformer\ntrained on WMT 2014 English-to-German, and ResNet-50 trained on ImageNet.\nAcross thousands of experiments, we demonstrate that complex techniques\n(Molchanov et al., 2017; Louizos et al., 2017b) shown to yield high compression\nrates on smaller datasets perform inconsistently, and that simple magnitude\npruning approaches achieve comparable or better results. Additionally, we\nreplicate the experiments performed by (Frankle & Carbin, 2018) and (Liu et\nal., 2018) at scale and show that unstructured sparse architectures learned\nthrough pruning cannot be trained from scratch to the same test set performance\nas a model trained with joint sparsification and optimization. Together, these\nresults highlight the need for large-scale benchmarks in the field of model\ncompression. We open-source our code, top performing model checkpoints, and\nresults of all hyperparameter configurations to establish rigorous baselines\nfor future work on compression and sparsification.","url_abs":"http://arxiv.org/abs/1902.09574v1","url_pdf":"http://arxiv.org/pdf/1902.09574v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-state-of-sparsity-in-deep-neural-networks","repo_url":"https://github.com/ars-ashuha/variational-dropout-sparsifies-dnn","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"GPL-3.0"}},{"paper_slug":"the-state-of-sparsity-in-deep-neural-networks","repo_url":"https://github.com/Shiweiliuiiiiiii/GraNet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"the-state-of-sparsity-in-deep-neural-networks","repo_url":"https://github.com/WilliamLiPro/LpSS","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"the-state-of-sparsity-in-deep-neural-networks","repo_url":"https://github.com/senya-ashukha/variational-dropout-sparsifies-dnn","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"GPL-3.0"}},{"paper_slug":"the-state-of-sparsity-in-deep-neural-networks","repo_url":"https://github.com/vita-group/granet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"the-state-of-sparsity-in-deep-neural-networks","repo_url":"https://github.com/google-research/google-research","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"model-compression","task_name":"Model Compression"},{"task_slug":"sparse-learning","task_name":"Sparse Learning"}],"methods":[{"method_slug":"pruning","method_name":"Pruning"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1902.09574","atlas_url":"https://app.syntology.ai/?focus=1902.09574","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}