{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/gasl-guided-attention-for-sparsity-learning","title":"GASL: Guided Attention for Sparsity Learning in Deep Neural Networks","arxiv_id":"1901.01939","date":"2019-01-07","proceeding":null,"authors":["Amirsina Torfi","Rouzbeh A. Shirvani","Sobhan Soleymani","Naser M. Nasrabadi"],"abstract":"The main goal of network pruning is imposing sparsity on the neural network\nby increasing the number of parameters with zero value in order to reduce the\narchitecture size and the computational speedup. In most of the previous\nresearch works, sparsity is imposed stochastically without considering any\nprior knowledge of the weights distribution or other internal network\ncharacteristics. Enforcing too much sparsity may induce accuracy drop due to\nthe fact that a lot of important elements might have been eliminated. In this\npaper, we propose Guided Attention for Sparsity Learning (GASL) to achieve (1)\nmodel compression by having less number of elements and speed-up; (2) prevent\nthe accuracy drop by supervising the sparsity operation via a guided attention\nmechanism and (3) introduce a generic mechanism that can be adapted for any\ntype of architecture; Our work is aimed at providing a framework based on\ninterpretable attention mechanisms for imposing structured and non-structured\nsparsity in deep neural networks. For Cifar-100 experiments, we achieved the\nstate-of-the-art sparsity level and 2.91x speedup with competitive accuracy\ncompared to the best method. For MNIST and LeNet architecture we also achieved\nthe highest sparsity and speedup level.","url_abs":"http://arxiv.org/abs/1901.01939v2","url_pdf":"http://arxiv.org/pdf/1901.01939v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"gasl-guided-attention-for-sparsity-learning","repo_url":"https://github.com/astorfi/attention-guided-sparsity","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"gasl-guided-attention-for-sparsity-learning","repo_url":"https://github.com/code-implementation1/Code4/tree/main/lenet","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"model-compression","task_name":"Model Compression"},{"task_slug":"network-pruning","task_name":"Network Pruning"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"lenet","method_name":"LeNet"},{"method_slug":"pruning","method_name":"Pruning"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1901.01939","atlas_url":"https://app.syntology.ai/?focus=1901.01939","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}