{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-evaluating-the-generalization-of-lstm","title":"On Evaluating the Generalization of LSTM Models in Formal Languages","arxiv_id":"1811.01001","date":"2018-11-02","proceeding":"WS 2019 1","authors":["Mirac Suzgun","Yonatan Belinkov","Stuart M. Shieber"],"abstract":"Recurrent Neural Networks (RNNs) are theoretically Turing-complete and\nestablished themselves as a dominant model for language processing. Yet, there\nstill remains an uncertainty regarding their language learning capabilities. In\nthis paper, we empirically evaluate the inductive learning capabilities of Long\nShort-Term Memory networks, a popular extension of simple RNNs, to learn simple\nformal languages, in particular $a^nb^n$, $a^nb^nc^n$, and $a^nb^nc^nd^n$. We\ninvestigate the influence of various aspects of learning, such as training data\nregimes and model capacity, on the generalization to unobserved samples. We\nfind striking differences in model performances under different training\nsettings and highlight the need for careful analysis and assessment when making\nclaims about the learning capabilities of neural network models.","url_abs":"http://arxiv.org/abs/1811.01001v1","url_pdf":"http://arxiv.org/pdf/1811.01001v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-evaluating-the-generalization-of-lstm","repo_url":"https://github.com/suzgunmirac/lstm-eval","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"inductive-learning","task_name":"Inductive Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1811.01001","atlas_url":"https://app.syntology.ai/?focus=1811.01001","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}