{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/universal-dependency-parsing-for-hindi","title":"Universal Dependency Parsing for Hindi-English Code-switching","arxiv_id":"1804.05868","date":"2018-04-16","proceeding":"NAACL 2018 6","authors":["Irshad Ahmad Bhat","Riyaz Ahmad Bhat","Manish Shrivastava","Dipti Misra Sharma"],"abstract":"Code-switching is a phenomenon of mixing grammatical structures of two or\nmore languages under varied social constraints. The code-switching data differ\nso radically from the benchmark corpora used in NLP community that the\napplication of standard technologies to these data degrades their performance\nsharply. Unlike standard corpora, these data often need to go through\nadditional processes such as language identification, normalization and/or\nback-transliteration for their efficient processing. In this paper, we\ninvestigate these indispensable processes and other problems associated with\nsyntactic parsing of code-switching data and propose methods to mitigate their\neffects. In particular, we study dependency parsing of code-switching data of\nHindi and English multilingual speakers from Twitter. We present a treebank of\nHindi-English code-switching tweets under Universal Dependencies scheme and\npropose a neural stacking model for parsing that efficiently leverages\npart-of-speech tag and syntactic tree annotations in the code-switching\ntreebank and the preexisting Hindi and English treebanks. We also present\nnormalization and back-transliteration models with a decoding process tailored\nfor code-switching data. Results show that our neural stacking parser is 1.5%\nLAS points better than the augmented parsing model and our decoding process\nimproves results by 3.8% LAS points over the first-best normalization and/or\nback-transliteration.","url_abs":"http://arxiv.org/abs/1804.05868v3","url_pdf":"http://arxiv.org/pdf/1804.05868v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"universal-dependency-parsing-for-hindi","repo_url":"https://github.com/CodeMixedUniversalDependencies/UD_Hindi_English","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null},{"paper_slug":"universal-dependency-parsing-for-hindi","repo_url":"https://github.com/irshadbhat/nsdp-cs","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"dependency-parsing","task_name":"Dependency Parsing"},{"task_slug":"language-identification","task_name":"Language Identification"},{"task_slug":"tag","task_name":"TAG"},{"task_slug":"transliteration","task_name":"Transliteration"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1804.05868","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}