{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/venom-a-vectorized-n-m-format-for-unleashing","title":"VENOM: A Vectorized N:M Format for Unleashing the Power of Sparse Tensor Cores","arxiv_id":"2310.02065","date":"2023-10-03","proceeding":null,"authors":["Roberto L. Castro","Andrei Ivanov","Diego Andrade","Tal Ben-Nun","Basilio B. Fraguela","Torsten Hoefler"],"abstract":"The increasing success and scaling of Deep Learning models demands higher computational efficiency and power. Sparsification can lead to both smaller models as well as higher compute efficiency, and accelerated hardware is becoming available. However, exploiting it efficiently requires kernel implementations, pruning algorithms, and storage formats, to utilize hardware support of specialized sparse vector units. An example of those are the NVIDIA's Sparse Tensor Cores (SPTCs), which promise a 2x speedup. However, SPTCs only support the 2:4 format, limiting achievable sparsity ratios to 50%. We present the V:N:M format, which enables the execution of arbitrary N:M ratios on SPTCs. To efficiently exploit the resulting format, we propose Spatha, a high-performance sparse-library for DL routines. We show that Spatha achieves up to 37x speedup over cuBLAS. We also demonstrate a second-order pruning technique that enables sparsification to high sparsity ratios with V:N:M and little to no loss in accuracy in modern transformers.","url_abs":"https://arxiv.org/abs/2310.02065v1","url_pdf":"https://arxiv.org/pdf/2310.02065v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"venom-a-vectorized-n-m-format-for-unleashing","repo_url":"https://github.com/udc-gac/venom","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"computational-efficiency","task_name":"Computational Efficiency"}],"methods":[{"method_slug":"pruning","method_name":"Pruning"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2310.02065","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2310.02065"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/udc-gac/venom","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran":3,"unverified":1},"by_repo_kind":{"official":{"samples":4,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"866bd271c87031f8","entry":"nmSparsifier","repo":"udc-gac/venom","repo_kind":"official","path":"benchmark/energy.py","file_url":"https://github.com/udc-gac/venom/blob/HEAD/benchmark/energy.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"866bd271c87031f8"}},{"code_sha256_prefix":"235f498ae4f93c36","entry":"round_up","repo":"udc-gac/venom","repo_kind":"official","path":"end2end/grouped_nmv_tensor.py","file_url":"https://github.com/udc-gac/venom/blob/HEAD/end2end/grouped_nmv_tensor.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"235f498ae4f93c36"}},{"code_sha256_prefix":"721bc89c148723fd","entry":"stringify","repo":"udc-gac/venom","repo_kind":"official","path":"benchmark/native_scripting.py","file_url":"https://github.com/udc-gac/venom/blob/HEAD/benchmark/native_scripting.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"721bc89c148723fd"}},{"code_sha256_prefix":"fea9da0420300a3a","entry":"compile","repo":"udc-gac/venom","repo_kind":"official","path":"benchmark/native_scripting.py","file_url":"https://github.com/udc-gac/venom/blob/HEAD/benchmark/native_scripting.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"fea9da0420300a3a"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}