{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-the-use-of-emojis-to-train-emotion","title":"On the Use of Emojis to Train Emotion Classifiers","arxiv_id":"1902.08906","date":"2019-02-24","proceeding":null,"authors":["Wegdan Hussien","Mahmoud Al-Ayyoub","Yahya Tashtoush","Mohammed Al-Kabi"],"abstract":"Nowadays, the automatic detection of emotions is employed by many\napplications in different fields like security informatics, e-learning, humor\ndetection, targeted advertising, etc. Many of these applications focus on\nsocial media and treat this problem as a classification problem, which requires\npreparing training data. The typical method for annotating the training data by\nhuman experts is considered time consuming, labor intensive and sometimes prone\nto error. Moreover, such an approach is not easily extensible to new\ndomains/languages since such extensions require annotating new training data.\nIn this study, we propose a distant supervised learning approach where the\ntraining sentences are automatically annotated based on the emojis they have.\nSuch training data would be very cheap to produce compared with the manually\ncreated training data, thus, much larger training data can be easily obtained.\nOn the other hand, this training data would naturally have lower quality as it\nmay contain some errors in the annotation. Nonetheless, we experimentally show\nthat training classifiers on cheap, large and possibly erroneous data annotated\nusing this approach leads to more accurate results compared with training the\nsame classifiers on the more expensive, much smaller and error-free manually\nannotated training data. Our experiments are conducted on an in-house dataset\nof emotional Arabic tweets and the classifiers we consider are: Support Vector\nMachine (SVM), Multinomial Naive Bayes (MNB) and Random Forest (RF). In\naddition to experimenting with single classifiers, we also consider using an\nensemble of classifiers. The results show that using an automatically annotated\ntraining data (that is only one order of magnitude larger than the manually\nannotated one) gives better results in almost all settings considered.","url_abs":"http://arxiv.org/abs/1902.08906v2","url_pdf":"http://arxiv.org/pdf/1902.08906v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-the-use-of-emojis-to-train-emotion","repo_url":"https://github.com/malayyoub/emojis-to-train-emotion-classifiers","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"humor-detection","task_name":"Humor Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}