{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/why-are-sensitive-functions-hard-for","title":"Why are Sensitive Functions Hard for Transformers?","arxiv_id":"2402.09963","date":"2024-02-15","proceeding":null,"authors":["Michael Hahn","Mark Rofin"],"abstract":"Empirical studies have identified a range of learnability biases and limitations of transformers, such as a persistent difficulty in learning to compute simple formal languages such as PARITY, and a bias towards low-degree functions. However, theoretical understanding remains limited, with existing expressiveness theory either overpredicting or underpredicting realistic learning abilities. We prove that, under the transformer architecture, the loss landscape is constrained by the input-space sensitivity: Transformers whose output is sensitive to many parts of the input string inhabit isolated points in parameter space, leading to a low-sensitivity bias in generalization. We show theoretically and empirically that this theory unifies a broad array of empirical observations about the learning abilities and biases of transformers, such as their generalization bias towards low sensitivity and low degree, and difficulty in length generalization for PARITY. This shows that understanding transformers' inductive biases requires studying not just their in-principle expressivity, but also their loss landscape.","url_abs":"https://arxiv.org/abs/2402.09963v4","url_pdf":"https://arxiv.org/pdf/2402.09963v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"why-are-sensitive-functions-hard-for","repo_url":"https://github.com/lacoco-lab/sensitivity-hardness","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"sensitivity","task_name":"Sensitivity"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2402.09963","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2402.09963"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lacoco-lab/sensitivity-hardness","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":4},"by_repo_kind":{"official":{"samples":4,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"7f7f4130d5f9842f","entry":"compute_loss","repo":"lacoco-lab/sensitivity-hardness","repo_kind":"official","path":"src/utils.py","file_url":"https://github.com/lacoco-lab/sensitivity-hardness/blob/HEAD/src/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7f7f4130d5f9842f"}},{"code_sha256_prefix":"8401af6d2fb17199","entry":"fast_average_sensitivity","repo":"lacoco-lab/sensitivity-hardness","repo_kind":"official","path":"src/utils.py","file_url":"https://github.com/lacoco-lab/sensitivity-hardness/blob/HEAD/src/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"8401af6d2fb17199"}},{"code_sha256_prefix":"adbaebf298f00696","entry":"get_input","repo":"lacoco-lab/sensitivity-hardness","repo_kind":"official","path":"src/measurements.py","file_url":"https://github.com/lacoco-lab/sensitivity-hardness/blob/HEAD/src/measurements.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"adbaebf298f00696"}},{"code_sha256_prefix":"93e646a036ade7e6","entry":"get_parities_for_sensitivity","repo":"lacoco-lab/sensitivity-hardness","repo_kind":"official","path":"src/utils.py","file_url":"https://github.com/lacoco-lab/sensitivity-hardness/blob/HEAD/src/utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"93e646a036ade7e6"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}