Papers › BERTweet: A pre-trained language model for English Tweets

BERTweet: A pre-trained language model for English Tweets

20 May 2020EMNLP 2020 11arXiv:2005.10200archive 2025-07-28

Dat Quoc Nguyen, Thanh Vu, Anh Tuan Nguyen

We present BERTweet, the first public large-scale pre-trained language model for English Tweets. Our BERTweet, having the same architecture as BERT-base (Devlin et al., 2019), is trained using the RoBERTa pre-training procedure (Liu et al., 2019). Experiments show that BERTweet outperforms strong baselines RoBERTa-base and XLM-R-base (Conneau et al., 2020), producing better performance results than the previous state-of-the-art models on three Tweet NLP tasks: Part-of-speech tagging, Named-entity recognition and text classification. We release BERTweet under the MIT License to facilitate future research and applications on Tweet data. Our BERTweet is available at https://github.com/VinAIResearch/BERTweet

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

VinAIResearch/BERTweet officialmentioned in papermentioned on GitHubpytorchMIT report
cardiffnlp/tweeteval mentioned on GitHub report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Language ModelingLanguage ModellingNamed Entity RecognitionNamed Entity Recognition (NER)Part-Of-Speech TaggingSentiment AnalysisText ClassificationXLM-Rnamed-entity-recognitiontext-classification

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Named Entity Recognition (NER) WNUT 2016 BERTweet F1 52.1 #7 of 7 Archive leaderboard report
Named Entity Recognition (NER) WNUT 2017 BERTweet F1 56.5 #9 of 23 Archive leaderboard report
Part-Of-Speech Tagging Ritter BERTweet Acc 90.1 #4 of 4 Archive leaderboard report
Part-Of-Speech Tagging Tweebank BERTweet Acc 95.2 #2 of 3 Archive leaderboard report
Sentiment Analysis TweetEval BERTweet ALL 67.9 #1 of 7 Archive leaderboard report
Sentiment Analysis TweetEval BERTweet Emoji 33.4 #1 of 7 Archive leaderboard report
Sentiment Analysis TweetEval BERTweet Emotion 79.3 #1 of 7 Archive leaderboard report
Sentiment Analysis TweetEval BERTweet Irony 82.1 #1 of 7 Archive leaderboard report
Sentiment Analysis TweetEval BERTweet Offensive 79.5 #1 of 7 Archive leaderboard report
Sentiment Analysis TweetEval BERTweet Sentiment 73.4 #1 of 7 Archive leaderboard report
Sentiment Analysis TweetEval BERTweet Stance 71.2 #1 of 7 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AdamAttentionAttention DropoutBERTDense ConnectionsDropoutLayer NormalizationLinear LayerLinear Warmup With Linear DecayMulti-Head AttentionResidual ConnectionRoBERTaSoftmaxWeight DecayWordPiece

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections