{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-semantics-based-measure-of-emoji-similarity","title":"A Semantics-Based Measure of Emoji Similarity","arxiv_id":"1707.04653","date":"2017-07-14","proceeding":null,"authors":["Sanjaya Wijeratne","Lakshika Balasuriya","Amit Sheth","Derek Doran"],"abstract":"Emoji have grown to become one of the most important forms of communication\non the web. With its widespread use, measuring the similarity of emoji has\nbecome an important problem for contemporary text processing since it lies at\nthe heart of sentiment analysis, search, and interface design tasks. This paper\npresents a comprehensive analysis of the semantic similarity of emoji through\nembedding models that are learned over machine-readable emoji meanings in the\nEmojiNet knowledge base. Using emoji descriptions, emoji sense labels and emoji\nsense definitions, and with different training corpora obtained from Twitter\nand Google News, we develop and test multiple embedding models to measure emoji\nsimilarity. To evaluate our work, we create a new dataset called EmoSim508,\nwhich assigns human-annotated semantic similarity scores to a set of 508\ncarefully selected emoji pairs. After validation with EmoSim508, we present a\nreal-world use-case of our emoji embedding models using a sentiment analysis\ntask and show that our models outperform the previous best-performing emoji\nembedding model on this task. The EmoSim508 dataset and our emoji embedding\nmodels are publicly released with this paper and can be downloaded from\nhttp://emojinet.knoesis.org/.","url_abs":"http://arxiv.org/abs/1707.04653v1","url_pdf":"http://arxiv.org/pdf/1707.04653v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-semantics-based-measure-of-emoji-similarity","repo_url":"https://github.com/hougrammer/emoji_project","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"a-semantics-based-measure-of-emoji-similarity","repo_url":"https://github.com/joonasrooben/NLP-text2emoji","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"semantic-similarity","task_name":"Semantic Similarity"},{"task_slug":"semantic-textual-similarity","task_name":"Semantic Textual Similarity"},{"task_slug":"sentiment-analysis","task_name":"Sentiment Analysis"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}