{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/enhance-word-representation-for-out-of","title":"Enhance word representation for out-of-vocabulary on Ubuntu dialogue corpus","arxiv_id":"1802.02614","date":"2018-02-07","proceeding":"ICLR 2018 1","authors":["Jianxiong Dong","Jim Huang"],"abstract":"Ubuntu dialogue corpus is the largest public available dialogue corpus to\nmake it feasible to build end-to-end deep neural network models directly from\nthe conversation data. One challenge of Ubuntu dialogue corpus is the large\nnumber of out-of-vocabulary words. In this paper we proposed a method which\ncombines the general pre-trained word embedding vectors with those generated on\nthe task-specific training set to address this issue. We integrated character\nembedding into Chen et al's Enhanced LSTM method (ESIM) and used it to evaluate\nthe effectiveness of our proposed method. For the task of next utterance\nselection, the proposed method has demonstrated a significant performance\nimprovement against original ESIM and the new model has achieved\nstate-of-the-art results on both Ubuntu dialogue corpus and Douban conversation\ncorpus. In addition, we investigated the performance impact of end-of-utterance\nand end-of-turn token tags.","url_abs":"http://arxiv.org/abs/1802.02614v2","url_pdf":"http://arxiv.org/pdf/1802.02614v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"enhance-word-representation-for-out-of","repo_url":"https://github.com/jdongca2003/next_utterance_selection","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[],"methods":[{"method_slug":"esim","method_name":"ESIM"},{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1802.02614","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}