{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/classical-structured-prediction-losses-for","title":"Classical Structured Prediction Losses for Sequence to Sequence Learning","arxiv_id":"1711.04956","date":"2017-11-14","proceeding":"NAACL 2018 6","authors":["Sergey Edunov","Myle Ott","Michael Auli","David Grangier","Marc'Aurelio Ranzato"],"abstract":"There has been much recent work on training neural attention models at the\nsequence-level using either reinforcement learning-style methods or by\noptimizing the beam. In this paper, we survey a range of classical objective\nfunctions that have been widely used to train linear models for structured\nprediction and apply them to neural sequence to sequence models. Our\nexperiments show that these losses can perform surprisingly well by slightly\noutperforming beam search optimization in a like for like setup. We also report\nnew state of the art results on both IWSLT'14 German-English translation as\nwell as Gigaword abstractive summarization. On the larger WMT'14 English-French\ntranslation task, sequence-level training achieves 41.5 BLEU which is on par\nwith the state of the art.","url_abs":"http://arxiv.org/abs/1711.04956v5","url_pdf":"http://arxiv.org/pdf/1711.04956v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"classical-structured-prediction-losses-for","repo_url":"https://github.com/pytorch/fairseq","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"abstractive-text-summarization","task_name":"Abstractive Text Summarization"},{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"prediction","task_name":"Prediction"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"structured-prediction","task_name":"Structured Prediction"},{"task_slug":"translation","task_name":"Translation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/machine-translation-on-iwslt2014-german","task":"Machine Translation","dataset":"IWSLT2014 German-English","model":"Minimum Risk Training [Edunov2017]","rank_in_archive_order":30,"of":34,"metrics":{"BLEU score":"32.84"},"uses_additional_data":false},{"leaderboard":"/sota/machine-translation-on-iwslt2015-german","task":"Machine Translation","dataset":"IWSLT2015 German-English","model":"ConvS2S+Risk","rank_in_archive_order":4,"of":15,"metrics":{"BLEU score":"32.93"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1711.04956","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}