{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/data-augmentation-via-dependency-tree-1","title":"Data Augmentation via Dependency Tree Morphing for Low-Resource Languages","arxiv_id":"1903.09460","date":"2019-03-22","proceeding":"EMNLP 2018 10","authors":["Gözde Gül Şahin","Mark Steedman"],"abstract":"Neural NLP systems achieve high scores in the presence of sizable training\ndataset. Lack of such datasets leads to poor system performances in the case\nlow-resource languages. We present two simple text augmentation techniques\nusing dependency trees, inspired from image processing. We crop sentences by\nremoving dependency links, and we rotate sentences by moving the tree fragments\naround the root. We apply these techniques to augment the training sets of\nlow-resource languages in Universal Dependencies project. We implement a\ncharacter-level sequence tagging model and evaluate the augmented datasets on\npart-of-speech tagging task. We show that crop and rotate provides improvements\nover the models trained with non-augmented data for majority of the languages,\nespecially for languages with rich case marking systems.","url_abs":"http://arxiv.org/abs/1903.09460v1","url_pdf":"http://arxiv.org/pdf/1903.09460v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"data-augmentation-via-dependency-tree-1","repo_url":"https://github.com/gozdesahin/crop-rotate-augment","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"data-augmentation-via-dependency-tree-1","repo_url":"https://github.com/makcedward/nlpaug","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"part-of-speech-tagging","task_name":"Part-Of-Speech Tagging"},{"task_slug":"text-augmentation","task_name":"Text Augmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1903.09460","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}