{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/smpost-parts-of-speech-tagger-for-code-mixed","title":"SMPOST: Parts of Speech Tagger for Code-Mixed Indic Social Media Text","arxiv_id":"1702.00167","date":"2017-02-01","proceeding":null,"authors":["Deepak Gupta","Shubham Tripathi","Asif Ekbal","Pushpak Bhattacharyya"],"abstract":"Use of social media has grown dramatically during the last few years. Users\nfollow informal languages in communicating through social media. The language\nof communication is often mixed in nature, where people transcribe their\nregional language with English and this technique is found to be extremely\npopular. Natural language processing (NLP) aims to infer the information from\nthese text where Part-of-Speech (PoS) tagging plays an important role in\ngetting the prosody of the written text. For the task of PoS tagging on\nCode-Mixed Indian Social Media Text, we develop a supervised system based on\nConditional Random Field classifier. In order to tackle the problem\neffectively, we have focused on extracting rich linguistic features. We\nparticipate in three different language pairs, ie. English-Hindi,\nEnglish-Bengali and English-Telugu on three different social media platforms,\nTwitter, Facebook & WhatsApp. The proposed system is able to successfully\nassign coarse as well as fine-grained PoS tag labels for a given a code-mixed\nsentence. Experiments show that our system is quite generic that shows\nencouraging performance levels on all the three language pairs in all the\ndomains.","url_abs":"http://arxiv.org/abs/1702.00167v2","url_pdf":"http://arxiv.org/pdf/1702.00167v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"smpost-parts-of-speech-tagger-for-code-mixed","repo_url":"https://github.com/stripathi08/pos_cmism","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"pos","task_name":"POS"},{"task_slug":"pos-tagging","task_name":"POS Tagging"},{"task_slug":"part-of-speech-tagging","task_name":"Part-Of-Speech Tagging"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"tag","task_name":"TAG"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1702.00167","atlas_url":"https://app.syntology.ai/?focus=1702.00167","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}