{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/from-softmax-to-sparsemax-a-sparse-model-of","title":"From Softmax to Sparsemax: A Sparse Model of Attention and Multi-Label Classification","arxiv_id":"1602.02068","date":"2016-02-05","proceeding":null,"authors":["André F. T. Martins","Ramón Fernandez Astudillo"],"abstract":"We propose sparsemax, a new activation function similar to the traditional\nsoftmax, but able to output sparse probabilities. After deriving its\nproperties, we show how its Jacobian can be efficiently computed, enabling its\nuse in a network trained with backpropagation. Then, we propose a new smooth\nand convex loss function which is the sparsemax analogue of the logistic loss.\nWe reveal an unexpected connection between this new loss and the Huber\nclassification loss. We obtain promising empirical results in multi-label\nclassification problems and in attention-based neural networks for natural\nlanguage inference. For the latter, we achieve a similar performance as the\ntraditional softmax, but with a selective, more compact, attention focus.","url_abs":"http://arxiv.org/abs/1602.02068v2","url_pdf":"http://arxiv.org/pdf/1602.02068v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"from-softmax-to-sparsemax-a-sparse-model-of","repo_url":"https://github.com/AndreasMadsen/course-02456-sparsemax","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"from-softmax-to-sparsemax-a-sparse-model-of","repo_url":"https://github.com/aced125/sparsemax","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"from-softmax-to-sparsemax-a-sparse-model-of","repo_url":"https://github.com/deep-spin/entmax","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"from-softmax-to-sparsemax-a-sparse-model-of","repo_url":"https://github.com/dhruvdcoder/sparse-structured-attention","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"from-softmax-to-sparsemax-a-sparse-model-of","repo_url":"https://github.com/kriskorrel/sparsemax-pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"from-softmax-to-sparsemax-a-sparse-model-of","repo_url":"https://github.com/qrfaction/keras-sparsemax","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"from-softmax-to-sparsemax-a-sparse-model-of","repo_url":"https://github.com/vene/sparse-structured-attention","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"from-softmax-to-sparsemax-a-sparse-model-of","repo_url":"https://github.com/weiwang2330/sparse-structured-attention","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"from-softmax-to-sparsemax-a-sparse-model-of","repo_url":"https://github.com/MindCode-4/code-12/tree/main/from-softmax-to-sparsemax-a-sparse","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"from-softmax-to-sparsemax-a-sparse-model-of","repo_url":"https://github.com/MindCode-4/code-7/tree/main/from-softmax-to-sparsemax-a-sparse","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"from-softmax-to-sparsemax-a-sparse-model-of","repo_url":"https://github.com/MindSpore-scientific-2/code-12/tree/main/from-softmax-to-sparsemax-a-sparse","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"multi-label-classification-2","task_name":"MUlTI-LABEL-ClASSIFICATION"},{"task_slug":"multi-label-classification","task_name":"Multi-Label Classification"},{"task_slug":"natural-language-inference","task_name":"Natural Language Inference"}],"methods":[{"method_slug":"sparsemax","method_name":"Sparsemax"}],"datasets_introduced":[],"methods_introduced":[{"slug":"sparsemax","name":"Sparsemax","full_name":"Sparsemax"}],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1602.02068","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1602.02068"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/weiwang2330/sparse-structured-attention","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/aced125/sparsemax","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/AndreasMadsen/course-02456-sparsemax","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindCode-4/code-7/tree/main/from-softmax-to-sparsemax-a-sparse","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindCode-4/code-12/tree/main/from-softmax-to-sparsemax-a-sparse","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/kriskorrel/sparsemax-pytorch","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/qrfaction/keras-sparsemax","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MindSpore-scientific-2/code-12/tree/main/from-softmax-to-sparsemax-a-sparse","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/dhruvdcoder/sparse-structured-attention","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/deep-spin/entmax","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/vene/sparse-structured-attention","reach":{"status":"unanswered"}}],"summary":{"ran_honours":4,"ran_draft_wrong":2,"ran_violates":2},"by_repo_kind":{"listed":{"samples":6,"ran":6,"repositories":3}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"b95cf081a20568e0","entry":"Rop","repo":"AndreasMadsen/course-02456-sparsemax","repo_kind":"listed","path":"tensorflow_python/sparsemax.py","file_url":"https://github.com/AndreasMadsen/course-02456-sparsemax/blob/HEAD/tensorflow_python/sparsemax.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b95cf081a20568e0"}},{"code_sha256_prefix":"e41b90069aeca727","entry":"entmax15","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"e41b90069aeca727"}},{"code_sha256_prefix":"8d31b8450757150f","entry":"forward","repo":"AndreasMadsen/course-02456-sparsemax","repo_kind":"listed","path":"tensorflow_python/sparsemax.py","file_url":"https://github.com/AndreasMadsen/course-02456-sparsemax/blob/HEAD/tensorflow_python/sparsemax.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8d31b8450757150f"}},{"code_sha256_prefix":"6503187ed9b4502e","entry":"jacobian","repo":"AndreasMadsen/course-02456-sparsemax","repo_kind":"listed","path":"tensorflow_python/sparsemax.py","file_url":"https://github.com/AndreasMadsen/course-02456-sparsemax/blob/HEAD/tensorflow_python/sparsemax.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6503187ed9b4502e"}},{"code_sha256_prefix":"3aead050bb96dbb8","entry":"project_simplex","repo":"weiwang2330/sparse-structured-attention","repo_kind":"listed","path":"pytorch/torchsparseattn/sparsemax.py","file_url":"https://github.com/weiwang2330/sparse-structured-attention/blob/HEAD/pytorch/torchsparseattn/sparsemax.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"3aead050bb96dbb8"}},{"code_sha256_prefix":"1f581f76688efd21","entry":"project_simplex","repo":"dhruvdcoder/sparse-structured-attention","repo_kind":"listed","path":"torchsparseattn/sparsemax.py","file_url":"https://github.com/dhruvdcoder/sparse-structured-attention/blob/HEAD/torchsparseattn/sparsemax.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"1f581f76688efd21"}},{"code_sha256_prefix":"df0db81c189b01e7","entry":"sparsemax","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"df0db81c189b01e7"}},{"code_sha256_prefix":"4bc5ce61eaa1f3ee","entry":"sparsemax_grad","repo":"weiwang2330/sparse-structured-attention","repo_kind":"listed","path":"pytorch/torchsparseattn/sparsemax.py","file_url":"https://github.com/weiwang2330/sparse-structured-attention/blob/HEAD/pytorch/torchsparseattn/sparsemax.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"4bc5ce61eaa1f3ee"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}