{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/stronger-baselines-for-trustable-results-in","title":"Stronger Baselines for Trustable Results in Neural Machine Translation","arxiv_id":"1706.09733","date":"2017-06-29","proceeding":"WS 2017 8","authors":["Michael Denkowski","Graham Neubig"],"abstract":"Interest in neural machine translation has grown rapidly as its effectiveness\nhas been demonstrated across language and data scenarios. New research\nregularly introduces architectural and algorithmic improvements that lead to\nsignificant gains over \"vanilla\" NMT implementations. However, these new\ntechniques are rarely evaluated in the context of previously published\ntechniques, specifically those that are widely used in state-of-theart\nproduction and shared-task systems. As a result, it is often difficult to\ndetermine whether improvements from research will carry over to systems\ndeployed for real-world use. In this work, we recommend three specific methods\nthat are relatively easy to implement and result in much stronger experimental\nsystems. Beyond reporting significantly higher BLEU scores, we conduct an\nin-depth analysis of where improvements originate and what inherent weaknesses\nof basic NMT models are being addressed. We then compare the relative gains\nafforded by several other techniques proposed in the literature when starting\nwith vanilla systems versus our stronger baselines, showing that experimental\nconclusions may change depending on the baseline chosen. This indicates that\nchoosing a strong baseline is crucial for reporting reliable experimental\nresults.","url_abs":"http://arxiv.org/abs/1706.09733v1","url_pdf":"http://arxiv.org/pdf/1706.09733v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"stronger-baselines-for-trustable-results-in","repo_url":"https://github.com/ijauregiCMCRC/ReWE_NMT","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"nmt","task_name":"NMT"},{"task_slug":"translation","task_name":"Translation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1706.09733","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}