{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-emergence-of-number-and-syntax-units-in","title":"The emergence of number and syntax units in LSTM language models","arxiv_id":"1903.07435","date":"2019-03-18","proceeding":"NAACL 2019 6","authors":["Yair Lakretz","German Kruszewski","Theo Desbordes","Dieuwke Hupkes","Stanislas Dehaene","Marco Baroni"],"abstract":"Recent work has shown that LSTMs trained on a generic language modeling\nobjective capture syntax-sensitive generalizations such as long-distance number\nagreement. We have however no mechanistic understanding of how they accomplish\nthis remarkable feat. Some have conjectured it depends on heuristics that do\nnot truly take hierarchical structure into account. We present here a detailed\nstudy of the inner mechanics of number tracking in LSTMs at the single neuron\nlevel. We discover that long-distance number information is largely managed by\ntwo `number units'. Importantly, the behaviour of these units is partially\ncontrolled by other units independently shown to track syntactic structure. We\nconclude that LSTMs are, to some extent, implementing genuinely syntactic\nprocessing mechanisms, paving the way to a more general understanding of\ngrammatical encoding in LSTMs.","url_abs":"http://arxiv.org/abs/1903.07435v2","url_pdf":"http://arxiv.org/pdf/1903.07435v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-emergence-of-number-and-syntax-units-in","repo_url":"https://github.com/FAIRNS/Number_and_syntax_units_in_LSTM_LMs","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1903.07435","atlas_url":"https://app.syntology.ai/?focus=1903.07435","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}