{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/token-level-and-sequence-level-loss-smoothing","title":"Token-level and sequence-level loss smoothing for RNN language models","arxiv_id":"1805.05062","date":"2018-05-14","proceeding":"ACL 2018 7","authors":["Maha Elbayad","Laurent Besacier","Jakob Verbeek"],"abstract":"Despite the effectiveness of recurrent neural network language models, their\nmaximum likelihood estimation suffers from two limitations. It treats all\nsentences that do not match the ground truth as equally poor, ignoring the\nstructure of the output space. Second, it suffers from \"exposure bias\": during\ntraining tokens are predicted given ground-truth sequences, while at test time\nprediction is conditioned on generated output sequences. To overcome these\nlimitations we build upon the recent reward augmented maximum likelihood\napproach \\ie sequence-level smoothing that encourages the model to predict\nsentences close to the ground truth according to a given performance metric. We\nextend this approach to token-level loss smoothing, and propose improvements to\nthe sequence-level smoothing approach. Our experiments on two different tasks,\nimage captioning and machine translation, show that token-level and\nsequence-level loss smoothing are complementary, and significantly improve\nresults.","url_abs":"http://arxiv.org/abs/1805.05062v1","url_pdf":"http://arxiv.org/pdf/1805.05062v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"token-level-and-sequence-level-loss-smoothing","repo_url":"https://github.com/elbayadm/seq2seq","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"translation","task_name":"Translation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.05062","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}