{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-sub-word-level-compositions-for","title":"Towards Sub-Word Level Compositions for Sentiment Analysis of Hindi-English Code Mixed Text","arxiv_id":"1611.00472","date":"2016-11-02","proceeding":"COLING 2016 12","authors":["Ameya Prabhu","Aditya Joshi","Manish Shrivastava","Vasudeva Varma"],"abstract":"Sentiment analysis (SA) using code-mixed data from social media has several\napplications in opinion mining ranging from customer satisfaction to social\ncampaign analysis in multilingual societies. Advances in this area are impeded\nby the lack of a suitable annotated dataset. We introduce a Hindi-English\n(Hi-En) code-mixed dataset for sentiment analysis and perform empirical\nanalysis comparing the suitability and performance of various state-of-the-art\nSA methods in social media.\n  In this paper, we introduce learning sub-word level representations in LSTM\n(Subword-LSTM) architecture instead of character-level or word-level\nrepresentations. This linguistic prior in our architecture enables us to learn\nthe information about sentiment value of important morphemes. This also seems\nto work well in highly noisy text containing misspellings as shown in our\nexperiments which is demonstrated in morpheme-level feature maps learned by our\nmodel. Also, we hypothesize that encoding this linguistic prior in the\nSubword-LSTM architecture leads to the superior performance. Our system attains\naccuracy 4-5% greater than traditional approaches on our dataset, and also\noutperforms the available system for sentiment analysis in Hi-En code-mixed\ntext by 18%.","url_abs":"http://arxiv.org/abs/1611.00472v1","url_pdf":"http://arxiv.org/pdf/1611.00472v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"towards-sub-word-level-compositions-for","repo_url":"https://github.com/DrImpossible/Sub-word-LSTM","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"towards-sub-word-level-compositions-for","repo_url":"https://github.com/Vidyapiratha/FYP","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"towards-sub-word-level-compositions-for","repo_url":"https://github.com/mankadronit/60DaysofUdacity-Challenge","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"opinion-mining","task_name":"Opinion Mining"},{"task_slug":"sentiment-analysis","task_name":"Sentiment Analysis"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1611.00472","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}