{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/examining-the-tip-of-the-iceberg-a-data-set","title":"Examining the Tip of the Iceberg: A Data Set for Idiom Translation","arxiv_id":"1802.04681","date":"2018-02-13","proceeding":"LREC 2018 5","authors":["Marzieh Fadaee","Arianna Bisazza","Christof Monz"],"abstract":"Neural Machine Translation (NMT) has been widely used in recent years with\nsignificant improvements for many language pairs. Although state-of-the-art NMT\nsystems are generating progressively better translations, idiom translation\nremains one of the open challenges in this field. Idioms, a category of\nmultiword expressions, are an interesting language phenomenon where the overall\nmeaning of the expression cannot be composed from the meanings of its parts. A\nfirst important challenge is the lack of dedicated data sets for learning and\nevaluating idiom translation. In this paper we address this problem by creating\nthe first large-scale data set for idiom translation. Our data set is\nautomatically extracted from a widely used German-English translation corpus\nand includes, for each language direction, a targeted evaluation set where all\nsentences contain idioms and a regular training corpus where sentences\nincluding idioms are marked. We release this data set and use it to perform\npreliminary NMT experiments as the first step towards better idiom translation.","url_abs":"http://arxiv.org/abs/1802.04681v1","url_pdf":"http://arxiv.org/pdf/1802.04681v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"examining-the-tip-of-the-iceberg-a-data-set","repo_url":"https://github.com/marziehf/IdiomTranslationDS","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"nmt","task_name":"NMT"},{"task_slug":"translation","task_name":"Translation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1802.04681","atlas_url":"https://app.syntology.ai/?focus=1802.04681","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}