{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/alternative-structures-for-character-level","title":"Alternative structures for character-level RNNs","arxiv_id":"1511.06303","date":"2015-11-19","proceeding":null,"authors":["Piotr Bojanowski","Armand Joulin","Tomas Mikolov"],"abstract":"Recurrent neural networks are convenient and efficient models for language\nmodeling. However, when applied on the level of characters instead of words,\nthey suffer from several problems. In order to successfully model long-term\ndependencies, the hidden representation needs to be large. This in turn implies\nhigher computational costs, which can become prohibitive in practice. We\npropose two alternative structural modifications to the classical RNN model.\nThe first one consists on conditioning the character level representation on\nthe previous word representation. The other one uses the character history to\ncondition the output probability. We evaluate the performance of the two\nproposed modifications on challenging, multi-lingual real world data.","url_abs":"http://arxiv.org/abs/1511.06303v2","url_pdf":"http://arxiv.org/pdf/1511.06303v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"alternative-structures-for-character-level","repo_url":"https://github.com/cloudmcloudyo/capstone","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1511.06303","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}