{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/accelerated-reinforcement-learning-for","title":"Accelerated Reinforcement Learning for Sentence Generation by Vocabulary Prediction","arxiv_id":"1809.01694","date":"2018-09-05","proceeding":"NAACL 2019 6","authors":["Kazuma Hashimoto","Yoshimasa Tsuruoka"],"abstract":"A major obstacle in reinforcement learning-based sentence generation is the\nlarge action space whose size is equal to the vocabulary size of the\ntarget-side language. To improve the efficiency of reinforcement learning, we\npresent a novel approach for reducing the action space based on dynamic\nvocabulary prediction. Our method first predicts a fixed-size small vocabulary\nfor each input to generate its target sentence. The input-specific vocabularies\nare then used at supervised and reinforcement learning steps, and also at test\ntime. In our experiments on six machine translation and two image captioning\ndatasets, our method achieves faster reinforcement learning ($\\sim$2.7x faster)\nwith less GPU memory ($\\sim$2.3x less) than the full-vocabulary counterpart.\nThe reinforcement learning with our method consistently leads to significant\nimprovement of BLEU scores, and the scores are equal to or better than those of\nbaselines using the full vocabularies, with faster decoding time ($\\sim$3x\nfaster) on CPUs.","url_abs":"http://arxiv.org/abs/1809.01694v2","url_pdf":"http://arxiv.org/pdf/1809.01694v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"accelerated-reinforcement-learning-for","repo_url":"https://github.com/hassyGo/NLG-RL","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":null,"task_name":"GPU"},{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"prediction","task_name":"Prediction"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1809.01694","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}