{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/handling-rare-items-in-data-to-text","title":"Handling Rare Items in Data-to-Text Generation","arxiv_id":null,"date":"2018-11-01","proceeding":"WS 2018 11","authors":["Anastasia Shimorina","Claire Gardent"],"abstract":"Neural approaches to data-to-text generation generally handle rare input items using either delexicalisation or a copy mechanism. We investigate the relative impact of these two methods on two datasets (E2E and WebNLG) and using two evaluation settings. We show (i) that rare items strongly impact performance; (ii) that combining delexicalisation and copying yields the strongest improvement; (iii) that copying underperforms for rare and unseen items and (iv) that the impact of these two mechanisms greatly varies depending on how the dataset is constructed and on how it is split into train, dev and test.","url_abs":"https://aclanthology.org/W18-6543","url_pdf":"https://aclanthology.org/W18-6543.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"handling-rare-items-in-data-to-text","repo_url":"https://gitlab.com/shimorina/inlg-2018","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"handling-rare-items-in-data-to-text","repo_url":"https://gitlab.com/shimorina/webnlg-dataset","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"data-to-text-generation","task_name":"Data-to-Text Generation"},{"task_slug":"kg-to-text","task_name":"KG-to-Text Generation"},{"task_slug":"text-generation","task_name":"Text Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/kg-to-text-generation-on-webnlg-2-0-1","task":"KG-to-Text Generation","dataset":"WebNLG 2.0 (Constrained)","model":"SOTA-NPT","rank_in_archive_order":9,"of":9,"metrics":{"BLEU":"48.0","METEOR":"36.0","ROUGE":"65.0"},"uses_additional_data":false},{"leaderboard":"/sota/kg-to-text-generation-on-webnlg-2-0","task":"KG-to-Text Generation","dataset":"WebNLG 2.0 (Unconstrained)","model":"SOTA-NPT","rank_in_archive_order":11,"of":13,"metrics":{"BLEU":"61","METEOR":"42","ROUGE":"71.0"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}