{"url":"/method/roberta","slug":"roberta","name":"RoBERTa","full_name":"RoBERTa","full_name_withheld":false,"description_markdown":"**RoBERTa** is an extension of [BERT](https://paperswithcode.com/method/bert) with changes to the pretraining procedure. The modifications include: \r\n\r\n- training the model longer, with bigger batches, over more data\r\n- removing the next sentence prediction objective\r\n- training on longer sequences\r\n- dynamically changing the masking pattern applied to the training data. The authors also collect a large new dataset ($\\text{CC-News}$) of comparable size to other privately used datasets, to better control for training set size effects","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/1907.11692v1","title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Transformers","url":"/methods/category/transformers","pwc_aliases":[]}],"n_papers_tagged":913,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"Rethinking the effects of data contamination in Code Intelligence","date":"2025-06-03","arxiv_id":"2506.02791","n_code_links":0,"syntology":null},{"paper":"/paper/evaluating-ai-capabilities-in-detecting","title":"Evaluating AI capabilities in detecting conspiracy theories on YouTube","date":"2025-05-29","arxiv_id":"2505.23570","n_code_links":1,"syntology":null},{"paper":null,"title":"Detection of Suicidal Risk on Social Media: A Hybrid Model","date":"2025-05-26","arxiv_id":"2505.23797","n_code_links":0,"syntology":null},{"paper":"/paper/hermes-dravidianlangtech-2025-sentiment","title":"Hermes@DravidianLangTech 2025: Sentiment Analysis of Dravidian Languages using XLM-RoBERTa","date":"2025-05-25","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/optimized-text-embedding-models-and","title":"Optimized Text Embedding Models and Benchmarks for Amharic Passage Retrieval","date":"2025-05-25","arxiv_id":"2505.19356","n_code_links":1,"syntology":null},{"paper":null,"title":"Multi-Scale Probabilistic Generation Theory: A Hierarchical Framework for Interpreting Large Language Models","date":"2025-05-23","arxiv_id":"2505.18244","n_code_links":0,"syntology":null},{"paper":"/paper/climate-research-domain-berts-pretraining","title":"Climate Research Domain BERTs: Pretraining, Adaptation, and Evaluation","date":"2025-05-19","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":null,"title":"Suicide Risk Assessment Using Multimodal Speech Features: A Study on the SW1 Challenge Dataset","date":"2025-05-19","arxiv_id":"2505.13069","n_code_links":0,"syntology":null},{"paper":"/paper/kgalign-joint-semantic-structural-knowledge","title":"KGAlign: Joint Semantic-Structural Knowledge Encoding for Multimodal Fake News Detection","date":"2025-05-18","arxiv_id":"2505.14714","n_code_links":1,"syntology":null},{"paper":null,"title":"Comparative sentiment analysis of public perception: Monkeypox vs. COVID-19 behavioral insights","date":"2025-05-12","arxiv_id":"2505.07430","n_code_links":0,"syntology":null},{"paper":null,"title":"The Sound of Populism: Distinct Linguistic Features Across Populist Variants","date":"2025-05-10","arxiv_id":"2505.07874","n_code_links":0,"syntology":null},{"paper":null,"title":"Leveraging Language Models for Automated Patient Record Linkage","date":"2025-04-21","arxiv_id":"2504.15261","n_code_links":0,"syntology":null},{"paper":null,"title":"LLMs as Data Annotators: How Close Are We to Human Performance","date":"2025-04-21","arxiv_id":"2504.15022","n_code_links":0,"syntology":null},{"paper":null,"title":"The Synthetic Imputation Approach: Generating Optimal Synthetic Texts For Underrepresented Categories In Supervised Classification Tasks","date":"2025-04-21","arxiv_id":"2504.15160","n_code_links":0,"syntology":null},{"paper":null,"title":"ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance","date":"2025-04-11","arxiv_id":"2504.08716","n_code_links":0,"syntology":null},{"paper":null,"title":"Beyond LLMs: A Linguistic Approach to Causal Graph Generation from Narrative Texts","date":"2025-04-10","arxiv_id":"2504.07459","n_code_links":0,"syntology":null},{"paper":null,"title":"Exploring Generative AI Techniques in Government: A Case Study","date":"2025-04-06","arxiv_id":"2504.10497","n_code_links":0,"syntology":null},{"paper":"/paper/a-thorough-benchmark-of-automatic-text","title":"A thorough benchmark of automatic text classification: From traditional approaches to large language models","date":"2025-04-02","arxiv_id":"2504.01930","n_code_links":1,"syntology":null},{"paper":"/paper/context-aware-toxicity-detection-in","title":"Context-Aware Toxicity Detection in Multiplayer Games: Integrating Domain-Adaptive Pretraining and Match Metadata","date":"2025-04-02","arxiv_id":"2504.01534","n_code_links":1,"syntology":null},{"paper":"/paper/transformer-based-named-entity-recognition-2","title":"Transformer-Based Named Entity Recognition for Automated Server Provisioning","date":"2025-04-01","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"Measuring Online Hate on 4chan using Pre-trained Deep Learning Models","date":"2025-03-30","arxiv_id":"2504.00045","n_code_links":0,"syntology":null},{"paper":null,"title":"Enhancing Recommender Systems Using Textual Embeddings from Pre-trained Language Models","date":"2025-03-24","arxiv_id":"2504.08746","n_code_links":0,"syntology":null},{"paper":null,"title":"LakotaBERT: A Transformer-based Model for Low Resource Lakota Language","date":"2025-03-23","arxiv_id":"2503.18212","n_code_links":0,"syntology":null},{"paper":null,"title":"Predicting Human Choice Between Textually Described Lotteries","date":"2025-03-18","arxiv_id":"2503.14004","n_code_links":0,"syntology":null},{"paper":null,"title":"An Evaluation of LLMs for Detecting Harmful Computing Terms","date":"2025-03-12","arxiv_id":"2503.09341","n_code_links":0,"syntology":null},{"paper":null,"title":"Multimodal Emotion Recognition and Sentiment Analysis in Multi-Party Conversation Contexts","date":"2025-03-09","arxiv_id":"2503.06805","n_code_links":0,"syntology":null},{"paper":null,"title":"Seeing Delta Parameters as JPEG Images: Data-Free Delta Compression with Discrete Cosine Transform","date":"2025-03-09","arxiv_id":"2503.06676","n_code_links":0,"syntology":null},{"paper":"/paper/constructions-are-revealed-in-word","title":"Constructions are Revealed in Word Distributions","date":"2025-03-08","arxiv_id":"2503.06048","n_code_links":1,"syntology":null},{"paper":null,"title":"Evaluating open-source Large Language Models for automated fact-checking","date":"2025-03-07","arxiv_id":"2503.05565","n_code_links":0,"syntology":null},{"paper":null,"title":"Intermediate-Task Transfer Learning: Leveraging Sarcasm Detection for Stance Detection","date":"2025-03-05","arxiv_id":"2503.03172","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/language-modelling","name":"Language Modelling","papers":145},{"task":"/task/language-modeling","name":"Language Modeling","papers":128},{"task":"/task/sentence","name":"Sentence","papers":103},{"task":"/task/sentiment-analysis","name":"Sentiment Analysis","papers":75},{"task":"/task/question-answering","name":"Question Answering","papers":68},{"task":"/task/text-classification","name":"Text Classification","papers":65},{"task":"/task/text-classification-1","name":"text-classification","papers":56},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":45},{"task":"/task/classification-1","name":"Classification","papers":44},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":43},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":42},{"task":"/task/named-entity-recognition-1","name":"Named Entity Recognition","papers":40},{"task":"/task/named-entity-recognition","name":"named-entity-recognition","papers":39},{"task":"/task/named-entity-recognition-ner","name":"Named Entity Recognition (NER)","papers":35},{"task":"/task/natural-language-inference","name":"Natural Language Inference","papers":33},{"task":"/task/cg","name":"NER","papers":32},{"task":"/task/reading-comprehension","name":"Reading Comprehension","papers":28},{"task":"/task/articles","name":"Articles","papers":26},{"task":"/task/word-embeddings","name":"Word Embeddings","papers":26},{"task":"/task/decoder","name":"Decoder","papers":23}],"tasks_shown":20,"n_tasks":477,"usage_by_year":[{"year":"2019","papers":19},{"year":"2020","papers":168},{"year":"2021","papers":205},{"year":"2022","papers":138},{"year":"2023","papers":172},{"year":"2024","papers":159},{"year":"2025","papers":52}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/roberta"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}