{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/long-short-term-memory-for-japanese-word","title":"Long Short-Term Memory for Japanese Word Segmentation","arxiv_id":"1709.08011","date":"2017-09-23","proceeding":"PACLIC 2018 12","authors":["Yoshiaki Kitagawa","Mamoru Komachi"],"abstract":"This study presents a Long Short-Term Memory (LSTM) neural network approach\nto Japanese word segmentation (JWS). Previous studies on Chinese word\nsegmentation (CWS) succeeded in using recurrent neural networks such as LSTM\nand gated recurrent units (GRU). However, in contrast to Chinese, Japanese\nincludes several character types, such as hiragana, katakana, and kanji, that\nproduce orthographic variations and increase the difficulty of word\nsegmentation. Additionally, it is important for JWS tasks to consider a global\ncontext, and yet traditional JWS approaches rely on local features. In order to\naddress this problem, this study proposes employing an LSTM-based approach to\nJWS. The experimental results indicate that the proposed model achieves\nstate-of-the-art accuracy with respect to various Japanese corpora.","url_abs":"http://arxiv.org/abs/1709.08011v3","url_pdf":"http://arxiv.org/pdf/1709.08011v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"chinese-word-segmentation","task_name":"Chinese Word Segmentation"},{"task_slug":"japanese-word-segmentation","task_name":"Japanese Word Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/japanese-word-segmentation-on-bccwj","task":"Japanese Word Segmentation","dataset":"BCCWJ","model":"LSTM","rank_in_archive_order":3,"of":3,"metrics":{"F1-score (Word)":"0.9842"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}