Papers › TweetEval: Unified Benchmark and Comparative Evaluation for Tweet Classification

TweetEval: Unified Benchmark and Comparative Evaluation for Tweet Classification

23 Oct 2020Findings of the Association for Computational Linguistics 2020arXiv:2010.12421archive 2025-07-28

Francesco Barbieri, Jose Camacho-Collados, Leonardo Neves, Luis Espinosa-Anke

The experimental landscape in natural language processing for social media is too fragmented. Each year, new shared tasks and datasets are proposed, ranging from classics like sentiment analysis to irony detection or emoji prediction. Therefore, it is unclear what the current state of the art is, as there is no standardized evaluation protocol, neither a strong set of baselines trained on such domain-specific data. In this paper, we propose a new evaluation framework (TweetEval) consisting of seven heterogeneous Twitter-specific classification tasks. We also provide a strong set of baselines as starting point, and compare different language modeling pre-training strategies. Our initial experiments show the effectiveness of starting off with existing pre-trained generic language models, and continue training them on Twitter corpora.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

cardiffnlp/tweeteval officialmentioned in papermentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ClassificationGeneral ClassificationLanguage ModelingLanguage ModellingSentiment Analysis

Datasets

Introduced by this paper, per the archive.

TweetEval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Sentiment Analysis TweetEval RoBERTa-Base ALL 61.3 #3 of 7 Archive leaderboard report
Sentiment Analysis TweetEval RoBERTa-Base Emoji 30.9 #3 of 7 Archive leaderboard report
Sentiment Analysis TweetEval RoBERTa-Base Emotion 76.1 #3 of 7 Archive leaderboard report
Sentiment Analysis TweetEval RoBERTa-Base Hate 46.6 #3 of 7 Archive leaderboard report
Sentiment Analysis TweetEval RoBERTa-Base Irony 59.7 #3 of 7 Archive leaderboard report
Sentiment Analysis TweetEval RoBERTa-Base Offensive 79.5 #3 of 7 Archive leaderboard report
Sentiment Analysis TweetEval RoBERTa-Base Sentiment 71.3 #3 of 7 Archive leaderboard report
Sentiment Analysis TweetEval RoBERTa-Base Stance 68 #3 of 7 Archive leaderboard report
Sentiment Analysis TweetEval RoBERTa-Twitter ALL 61.0 #4 of 7 Archive leaderboard report
Sentiment Analysis TweetEval RoBERTa-Twitter Emoji 29.3 #4 of 7 Archive leaderboard report
Sentiment Analysis TweetEval RoBERTa-Twitter Emotion 72.0 #4 of 7 Archive leaderboard report
Sentiment Analysis TweetEval RoBERTa-Twitter Hate 49.9 #4 of 7 Archive leaderboard report
Sentiment Analysis TweetEval RoBERTa-Twitter Irony 65.4 #4 of 7 Archive leaderboard report
Sentiment Analysis TweetEval RoBERTa-Twitter Offensive 77.1 #4 of 7 Archive leaderboard report
Sentiment Analysis TweetEval RoBERTa-Twitter Sentiment 69.1 #4 of 7 Archive leaderboard report
Sentiment Analysis TweetEval RoBERTa-Twitter Stance 66.7 #4 of 7 Archive leaderboard report
Sentiment Analysis TweetEval SVM ALL 53.5 #5 of 7 Archive leaderboard report
Sentiment Analysis TweetEval SVM Emoji 29.3 #5 of 7 Archive leaderboard report
Sentiment Analysis TweetEval SVM Emotion 64.7 #5 of 7 Archive leaderboard report
Sentiment Analysis TweetEval SVM Hate 36.7 #5 of 7 Archive leaderboard report
Sentiment Analysis TweetEval SVM Irony 61.7 #5 of 7 Archive leaderboard report
Sentiment Analysis TweetEval SVM Offensive 52.3 #5 of 7 Archive leaderboard report
Sentiment Analysis TweetEval SVM Sentiment 62.9 #5 of 7 Archive leaderboard report
Sentiment Analysis TweetEval SVM Stance 67.3 #5 of 7 Archive leaderboard report
Sentiment Analysis TweetEval FastText ALL 58.1 #6 of 7 Archive leaderboard report
Sentiment Analysis TweetEval FastText Emoji 25.8 #6 of 7 Archive leaderboard report
Sentiment Analysis TweetEval FastText Emotion 65.2 #6 of 7 Archive leaderboard report
Sentiment Analysis TweetEval FastText Hate 50.6 #6 of 7 Archive leaderboard report
Sentiment Analysis TweetEval FastText Irony 63.1 #6 of 7 Archive leaderboard report
Sentiment Analysis TweetEval FastText Offensive 73.4 #6 of 7 Archive leaderboard report
Sentiment Analysis TweetEval FastText Sentiment 62.9 #6 of 7 Archive leaderboard report
Sentiment Analysis TweetEval FastText Stance 65.4 #6 of 7 Archive leaderboard report
Sentiment Analysis TweetEval LSTM ALL 56.5 #7 of 7 Archive leaderboard report
Sentiment Analysis TweetEval LSTM Emoji 24.7 #7 of 7 Archive leaderboard report
Sentiment Analysis TweetEval LSTM Emotion 66.0 #7 of 7 Archive leaderboard report
Sentiment Analysis TweetEval LSTM Hate 52.6 #7 of 7 Archive leaderboard report
Sentiment Analysis TweetEval LSTM Irony 62.8 #7 of 7 Archive leaderboard report
Sentiment Analysis TweetEval LSTM Offensive 71.7 #7 of 7 Archive leaderboard report
Sentiment Analysis TweetEval LSTM Sentiment 58.3 #7 of 7 Archive leaderboard report
Sentiment Analysis TweetEval LSTM Stance 59.4 #7 of 7 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections