{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-corpus-of-english-hindi-code-mixed-tweets","title":"A Corpus of English-Hindi Code-Mixed Tweets for Sarcasm Detection","arxiv_id":"1805.11869","date":"2018-05-30","proceeding":null,"authors":["Sahil Swami","Ankush Khandelwal","Vinay Singh","Syed Sarfaraz Akhtar","Manish Shrivastava"],"abstract":"Social media platforms like twitter and facebook have be- come two of the\nlargest mediums used by people to express their views to- wards different\ntopics. Generation of such large user data has made NLP tasks like sentiment\nanalysis and opinion mining much more important. Using sarcasm in texts on\nsocial media has become a popular trend lately. Using sarcasm reverses the\nmeaning and polarity of what is implied by the text which poses challenge for\nmany NLP tasks. The task of sarcasm detection in text is gaining more and more\nimportance for both commer- cial and security services. We present the first\nEnglish-Hindi code-mixed dataset of tweets marked for presence of sarcasm and\nirony where each token is also annotated with a language tag. We present a\nbaseline su- pervised classification system developed using the same dataset\nwhich achieves an average F-score of 78.4 after using random forest classifier\nand performing 10-fold cross validation.","url_abs":"http://arxiv.org/abs/1805.11869v1","url_pdf":"http://arxiv.org/pdf/1805.11869v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-corpus-of-english-hindi-code-mixed-tweets","repo_url":"https://github.com/asking28/offenseval2020","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"a-corpus-of-english-hindi-code-mixed-tweets","repo_url":"https://github.com/asking28/sentimix2020","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"opinion-mining","task_name":"Opinion Mining"},{"task_slug":"sarcasm-detection","task_name":"Sarcasm Detection"},{"task_slug":"sentiment-analysis","task_name":"Sentiment Analysis"},{"task_slug":"tag","task_name":"TAG"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1805.11869","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}