{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2607-21752","title":"Parameter-free Adaptive Sparse Attention via Compression-Based Content Selection","arxiv_id":"2607.21752","date":"2026-07-23","proceeding":null,"authors":["Debarshi Kundu","Swaroop Ghosh","Vasant Honavar"],"abstract":"Data-adaptive sparse attention masks substantially outperform fixed patterns (e.g., BigBird and Longformer) and can even exceed dense attention on long sequences. Existing adaptive approaches---including SBM-Transformer, Dynamic Mask Attention, and NSA---typically require additional learnable parameters, custom gradient estimators, or specialized CUDA kernels. We show that classical data compression provides an effective masking signal with \\textbf{no additional parameters}. By computing per-block gzip compression ratios, we identify non-redundant content blocks and route long-range attention selectively through them. Intuitively, blocks that gzip cannot compress contain information not predictable from local repetition, making them natural long-range attention targets. Because the compression profile is input-dependent, the resulting sparse mask adapts dynamically to content without learned parameters, auxiliary losses, or custom kernels. On PG-19 byte-level language modeling at 92M parameters with 8K context, our method achieves 1.71 bits-per-byte (BPB), outperforming dense attention (2.89), BigBird (2.34), Longformer (3.21), and a reimplemented SBM-Transformer (3.38)---the only learned-mask baseline---by up to 1.67 BPB while adding no parameters. The advantage grows with sequence length, with the gap over BigBird widening from 0.05 BPB at 4K context to 0.63 BPB at 8K, while convergence is 3.3$\\times$ faster.","url_abs":"https://arxiv.org/abs/2607.21752","url_pdf":"https://arxiv.org/pdf/2607.21752","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2607.21752","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2607.21752"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/sc782/SBM-Transformer","reach":null}],"summary":{"ran":7},"by_repo_kind":{"found_in_text":{"samples":7,"ran":7,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":7,"samples":[{"code_sha256_prefix":"defcc1e5e9bf48b2","entry":"append_cls","repo":"sc782/SBM-Transformer","repo_kind":"found_in_text","path":"code/model_wrapper.py","file_url":"https://github.com/sc782/SBM-Transformer/blob/HEAD/code/model_wrapper.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"defcc1e5e9bf48b2"}},{"code_sha256_prefix":"9852cec20a1d486e","entry":"attach_dim","repo":"sc782/SBM-Transformer","repo_kind":"found_in_text","path":"code/attention_sbm.py","file_url":"https://github.com/sc782/SBM-Transformer/blob/HEAD/code/attention_sbm.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9852cec20a1d486e"}},{"code_sha256_prefix":"5f1d8d3803a6d549","entry":"batched_bincount","repo":"sc782/SBM-Transformer","repo_kind":"found_in_text","path":"code/fastRG.py","file_url":"https://github.com/sc782/SBM-Transformer/blob/HEAD/code/fastRG.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5f1d8d3803a6d549"}},{"code_sha256_prefix":"04d8eb683eb1ce24","entry":"batched_multinomial","repo":"sc782/SBM-Transformer","repo_kind":"found_in_text","path":"code/fastRG.py","file_url":"https://github.com/sc782/SBM-Transformer/blob/HEAD/code/fastRG.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"04d8eb683eb1ce24"}},{"code_sha256_prefix":"2db10d7fc567bec3","entry":"block_diag","repo":"sc782/SBM-Transformer","repo_kind":"found_in_text","path":"code/attention_sbm.py","file_url":"https://github.com/sc782/SBM-Transformer/blob/HEAD/code/attention_sbm.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"2db10d7fc567bec3"}},{"code_sha256_prefix":"b1d7a0eee487a64b","entry":"fastRG","repo":"sc782/SBM-Transformer","repo_kind":"found_in_text","path":"code/fastRG.py","file_url":"https://github.com/sc782/SBM-Transformer/blob/HEAD/code/fastRG.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b1d7a0eee487a64b"}},{"code_sha256_prefix":"9dbe45c9859b04e5","entry":"pooling","repo":"sc782/SBM-Transformer","repo_kind":"found_in_text","path":"code/model_wrapper.py","file_url":"https://github.com/sc782/SBM-Transformer/blob/HEAD/code/model_wrapper.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9dbe45c9859b04e5"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}