{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/tweeteval-unified-benchmark-and-comparative","title":"TweetEval: Unified Benchmark and Comparative Evaluation for Tweet Classification","arxiv_id":"2010.12421","date":"2020-10-23","proceeding":"Findings of the Association for Computational Linguistics 2020","authors":["Francesco Barbieri","Jose Camacho-Collados","Leonardo Neves","Luis Espinosa-Anke"],"abstract":"The experimental landscape in natural language processing for social media is too fragmented. Each year, new shared tasks and datasets are proposed, ranging from classics like sentiment analysis to irony detection or emoji prediction. Therefore, it is unclear what the current state of the art is, as there is no standardized evaluation protocol, neither a strong set of baselines trained on such domain-specific data. In this paper, we propose a new evaluation framework (TweetEval) consisting of seven heterogeneous Twitter-specific classification tasks. We also provide a strong set of baselines as starting point, and compare different language modeling pre-training strategies. Our initial experiments show the effectiveness of starting off with existing pre-trained generic language models, and continue training them on Twitter corpora.","url_abs":"https://arxiv.org/abs/2010.12421v2","url_pdf":"https://arxiv.org/pdf/2010.12421v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"tweeteval-unified-benchmark-and-comparative","repo_url":"https://github.com/cardiffnlp/tweeteval","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}},{"paper_slug":"tweeteval-unified-benchmark-and-comparative","repo_url":"https://github.com/jinhxu/how-much-hate-with-china","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"sentiment-analysis","task_name":"Sentiment Analysis"}],"methods":[],"datasets_introduced":[{"slug":"tweeteval","name":"TweetEval","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/sentiment-analysis-on-tweeteval","task":"Sentiment Analysis","dataset":"TweetEval","model":"RoBERTa-Base","rank_in_archive_order":3,"of":7,"metrics":{"ALL":"61.3","Emoji":"30.9","Emotion":"76.1","Hate":"46.6","Irony":"59.7","Offensive":"79.5","Sentiment":"71.3","Stance":"68"},"uses_additional_data":false},{"leaderboard":"/sota/sentiment-analysis-on-tweeteval","task":"Sentiment Analysis","dataset":"TweetEval","model":"RoBERTa-Twitter","rank_in_archive_order":4,"of":7,"metrics":{"ALL":"61.0","Emoji":"29.3","Emotion":"72.0","Hate":"49.9","Irony":"65.4","Offensive":"77.1","Sentiment":"69.1","Stance":"66.7"},"uses_additional_data":false},{"leaderboard":"/sota/sentiment-analysis-on-tweeteval","task":"Sentiment Analysis","dataset":"TweetEval","model":"SVM","rank_in_archive_order":5,"of":7,"metrics":{"ALL":"53.5","Emoji":"29.3","Emotion":"64.7","Hate":"36.7","Irony":"61.7","Offensive":"52.3","Sentiment":"62.9","Stance":"67.3"},"uses_additional_data":false},{"leaderboard":"/sota/sentiment-analysis-on-tweeteval","task":"Sentiment Analysis","dataset":"TweetEval","model":"FastText","rank_in_archive_order":6,"of":7,"metrics":{"ALL":"58.1","Emoji":"25.8","Emotion":"65.2","Hate":"50.6","Irony":"63.1","Offensive":"73.4","Sentiment":"62.9","Stance":"65.4"},"uses_additional_data":false},{"leaderboard":"/sota/sentiment-analysis-on-tweeteval","task":"Sentiment Analysis","dataset":"TweetEval","model":"LSTM","rank_in_archive_order":7,"of":7,"metrics":{"ALL":"56.5","Emoji":"24.7","Emotion":"66.0","Hate":"52.6","Irony":"62.8","Offensive":"71.7","Sentiment":"58.3","Stance":"59.4"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2010.12421","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}