{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sequence-to-sequence-data-augmentation-for","title":"Sequence-to-Sequence Data Augmentation for Dialogue Language Understanding","arxiv_id":"1807.01554","date":"2018-07-04","proceeding":"COLING 2018 8","authors":["Yutai Hou","Yijia Liu","Wanxiang Che","Ting Liu"],"abstract":"In this paper, we study the problem of data augmentation for language\nunderstanding in task-oriented dialogue system. In contrast to previous work\nwhich augments an utterance without considering its relation with other\nutterances, we propose a sequence-to-sequence generation based data\naugmentation framework that leverages one utterance's same semantic\nalternatives in the training data. A novel diversity rank is incorporated into\nthe utterance representation to make the model produce diverse utterances and\nthese diversely augmented utterances help to improve the language understanding\nmodule. Experimental results on the Airline Travel Information System dataset\nand a newly created semantic frame annotation on Stanford Multi-turn,\nMultidomain Dialogue Dataset show that our framework achieves significant\nimprovements of 6.38 and 10.04 F-scores respectively when only a training set\nof hundreds utterances is represented. Case studies also confirm that our\nmethod generates diverse utterances.","url_abs":"http://arxiv.org/abs/1807.01554v1","url_pdf":"http://arxiv.org/pdf/1807.01554v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"sequence-to-sequence-data-augmentation-for","repo_url":"https://github.com/AtmaHou/Seq2SeqDataAugmentationForLU","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"text-augmentation","task_name":"Text Augmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1807.01554","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}