{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sparsifiner-learning-sparse-instance","title":"Sparsifiner: Learning Sparse Instance-Dependent Attention for Efficient Vision Transformers","arxiv_id":"2303.13755","date":"2023-03-24","proceeding":"CVPR 2023 1","authors":["Cong Wei","Brendan Duke","Ruowei Jiang","Parham Aarabi","Graham W. Taylor","Florian Shkurti"],"abstract":"Vision Transformers (ViT) have shown their competitive advantages performance-wise compared to convolutional neural networks (CNNs) though they often come with high computational costs. To this end, previous methods explore different attention patterns by limiting a fixed number of spatially nearby tokens to accelerate the ViT's multi-head self-attention (MHSA) operations. However, such structured attention patterns limit the token-to-token connections to their spatial relevance, which disregards learned semantic connections from a full attention mask. In this work, we propose a novel approach to learn instance-dependent attention patterns, by devising a lightweight connectivity predictor module to estimate the connectivity score of each pair of tokens. Intuitively, two tokens have high connectivity scores if the features are considered relevant either spatially or semantically. As each token only attends to a small number of other tokens, the binarized connectivity masks are often very sparse by nature and therefore provide the opportunity to accelerate the network via sparse computations. Equipped with the learned unstructured attention pattern, sparse attention ViT (Sparsifiner) produces a superior Pareto-optimal trade-off between FLOPs and top-1 accuracy on ImageNet compared to token sparsity. Our method reduces 48% to 69% FLOPs of MHSA while the accuracy drop is within 0.4%. We also show that combining attention and token sparsity reduces ViT FLOPs by over 60%.","url_abs":"https://arxiv.org/abs/2303.13755v1","url_pdf":"https://arxiv.org/pdf/2303.13755v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sparsifiner-learning-sparse-instance","repo_url":"https://github.com/lim142857/Sparsifiner","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2303.13755","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2303.13755"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lim142857/Sparsifiner","reach":null}],"summary":{"ran":1,"ran_draft_wrong":1,"unverified":3},"by_repo_kind":{"listed":{"samples":5,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"a0e2d527bd3523a6","entry":"MaskPredictor","repo":"lim142857/Sparsifiner","repo_kind":"listed","path":"src/models/sparsifiner.py","file_url":"https://github.com/lim142857/Sparsifiner/blob/HEAD/src/models/sparsifiner.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a0e2d527bd3523a6"}},{"code_sha256_prefix":"d6e08fa1f58591c6","entry":"compute_sparsity","repo":"lim142857/Sparsifiner","repo_kind":"listed","path":"src/models/sparsifiner.py","file_url":"https://github.com/lim142857/Sparsifiner/blob/HEAD/src/models/sparsifiner.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d6e08fa1f58591c6"}},{"code_sha256_prefix":"cd7c6fd7feccfdc2","entry":"Attention","repo":"lim142857/Sparsifiner","repo_kind":"listed","path":"src/models/sparsifiner.py","file_url":"https://github.com/lim142857/Sparsifiner/blob/HEAD/src/models/sparsifiner.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"cd7c6fd7feccfdc2"}},{"code_sha256_prefix":"6b461976af475942","entry":"Block","repo":"lim142857/Sparsifiner","repo_kind":"listed","path":"src/models/sparsifiner.py","file_url":"https://github.com/lim142857/Sparsifiner/blob/HEAD/src/models/sparsifiner.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6b461976af475942"}},{"code_sha256_prefix":"c7af31c985044e53","entry":"Sparsifiner","repo":"lim142857/Sparsifiner","repo_kind":"listed","path":"src/models/sparsifiner.py","file_url":"https://github.com/lim142857/Sparsifiner/blob/HEAD/src/models/sparsifiner.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c7af31c985044e53"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}