{"url":"/method/electra","slug":"electra","name":"ELECTRA","full_name":"ELECTRA","full_name_withheld":false,"description_markdown":"**ELECTRA** is a [transformer](https://paperswithcode.com/method/transformer) with a new pre-training approach which trains two transformer models: the generator and the discriminator. The generator replaces tokens in the sequence - trained as a masked language model - and the discriminator (the ELECTRA contribution) attempts to identify which tokens are replaced by the generator in the sequence. This pre-training task is called replaced token detection, and is a replacement for masking the input.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2003.10555v1","title":"ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Transformers","url":"/methods/category/transformers","pwc_aliases":[]}],"n_papers_tagged":118,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"AdUE: Improving uncertainty estimation head for LoRA adapters in LLMs","date":"2025-05-21","arxiv_id":"2505.15443","n_code_links":0,"syntology":null},{"paper":"/paper/litelmguard-seamless-and-lightweight-on","title":"LiteLMGuard: Seamless and Lightweight On-Device Prompt Filtering for Safeguarding Small Language Models against Quantization-induced Risks and Vulnerabilities","date":"2025-05-08","arxiv_id":"2505.05619","n_code_links":2,"syntology":null},{"paper":"/paper/sigma-a-dataset-for-text-to-code-semantic","title":"Sigma: A dataset for text-to-code semantic parsing with statistical analysis","date":"2025-04-05","arxiv_id":"2504.04301","n_code_links":1,"syntology":null},{"paper":null,"title":"Comparative Analysis of Efficient Adapter-Based Fine-Tuning of State-of-the-Art Transformer Models","date":"2025-01-14","arxiv_id":"2501.08271","n_code_links":0,"syntology":null},{"paper":"/paper/punctuation-s-semantic-role-between-brain-and","title":"Aligning Brain Activity with Advanced Transformer Models: Exploring the Role of Punctuation in Semantic Processing","date":"2025-01-10","arxiv_id":"2501.06278","n_code_links":1,"syntology":null},{"paper":null,"title":"IntegrityAI at GenAI Detection Task 2: Detecting Machine-Generated Academic Essays in English and Arabic Using ELECTRA and Stylometry","date":"2025-01-07","arxiv_id":"2501.05476","n_code_links":0,"syntology":null},{"paper":"/paper/electra-and-gpt-4o-cost-effective-partners","title":"ELECTRA and GPT-4o: Cost-Effective Partners for Sentiment Analysis","date":"2024-12-29","arxiv_id":"2501.00062","n_code_links":1,"syntology":null},{"paper":"/paper/what-differentiates-educational-literature-a","title":"What Differentiates Educational Literature? A Multimodal Fusion Approach of Transformers and Computational Linguistics","date":"2024-11-26","arxiv_id":"2411.17593","n_code_links":0,"syntology":null},{"paper":"/paper/a-comparative-analysis-of-transformer-and","title":"A Comparative Analysis of Transformer and LSTM Models for Detecting Suicidal Ideation on Reddit","date":"2024-11-23","arxiv_id":"2411.15404","n_code_links":1,"syntology":null},{"paper":null,"title":"LA4SR: illuminating the dark proteome with generative AI","date":"2024-11-11","arxiv_id":"2411.06798","n_code_links":0,"syntology":null},{"paper":null,"title":"A Named Entity Recognition and Topic Modeling-based Solution for Locating and Better Assessment of Natural Disasters in Social Media","date":"2024-05-01","arxiv_id":"2405.00903","n_code_links":0,"syntology":null},{"paper":"/paper/do-english-named-entity-recognizers-work-well","title":"Do \"English\" Named Entity Recognizers Work Well on Global Englishes?","date":"2024-04-20","arxiv_id":"2404.13465","n_code_links":2,"syntology":null},{"paper":"/paper/fastlogad-log-anomaly-detection-with-mask","title":"FastLogAD: Log Anomaly Detection with Mask-Guided Pseudo Anomaly Generation and Discrimination","date":"2024-04-12","arxiv_id":"2404.08750","n_code_links":1,"syntology":null},{"paper":"/paper/leveraging-pre-trained-language-models-for-3","title":"Leveraging pre-trained language models for code generation","date":"2024-02-29","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/are-electra-s-sentence-embeddings-beyond","title":"Are ELECTRA's Sentence Embeddings Beyond Repair? The Case of Semantic Textual Similarity","date":"2024-02-20","arxiv_id":"2402.13130","n_code_links":1,"syntology":null},{"paper":"/paper/a-curious-case-of-searching-for-the","title":"A Curious Case of Searching for the Correlation between Training Data and Adversarial Robustness of Transformer Textual Models","date":"2024-02-18","arxiv_id":"2402.11469","n_code_links":1,"syntology":null},{"paper":null,"title":"LogELECTRA: Self-supervised Anomaly Detection for Unstructured Logs","date":"2024-02-16","arxiv_id":"2402.10397","n_code_links":0,"syntology":null},{"paper":"/paper/dissecting-vocabulary-biases-datasets-through","title":"Dissecting vocabulary biases datasets through statistical testing and automated data augmentation for artifact mitigation in Natural Language Inference","date":"2023-12-14","arxiv_id":"2312.08747","n_code_links":1,"syntology":null},{"paper":null,"title":"FA Team at the NTCIR-17 UFO Task","date":"2023-10-31","arxiv_id":"2310.20322","n_code_links":0,"syntology":null},{"paper":null,"title":"Fast-ELECTRA for Efficient Pre-training","date":"2023-10-11","arxiv_id":"2310.07347","n_code_links":0,"syntology":null},{"paper":"/paper/structural-self-supervised-objectives-for","title":"Structural Self-Supervised Objectives for Transformers","date":"2023-09-15","arxiv_id":"2309.08272","n_code_links":1,"syntology":null},{"paper":"/paper/transformer-based-punctuation-restoration-for","title":"Transformer Based Punctuation Restoration for Turkish","date":"2023-09-15","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"Optimizing Multi-Class Text Classification: A Diverse Stacking Ensemble Framework Utilizing Transformers","date":"2023-08-19","arxiv_id":"2308.11519","n_code_links":0,"syntology":null},{"paper":null,"title":"An Ensemble Approach to Question Classification: Integrating Electra Transformer, GloVe, and LSTM","date":"2023-08-13","arxiv_id":"2308.06828","n_code_links":0,"syntology":null},{"paper":"/paper/analysis-of-the-evolution-of-advanced","title":"Analysis of the Evolution of Advanced Transformer-Based Language Models: Experiments on Opinion Mining","date":"2023-08-07","arxiv_id":"2308.03235","n_code_links":1,"syntology":null},{"paper":"/paper/identifying-misinformation-on-youtube-through","title":"Identifying Misinformation on YouTube through Transcript Contextual Analysis with Transformer Models","date":"2023-07-22","arxiv_id":"2307.12155","n_code_links":1,"syntology":null},{"paper":null,"title":"Vacaspati: A Diverse Corpus of Bangla Literature","date":"2023-07-11","arxiv_id":"2307.05083","n_code_links":0,"syntology":null},{"paper":null,"title":"Comparison of Pre-trained Language Models for Turkish Address Parsing","date":"2023-06-24","arxiv_id":"2306.13947","n_code_links":0,"syntology":null},{"paper":null,"title":"Named entity recognition in resumes","date":"2023-06-22","arxiv_id":"2306.13062","n_code_links":0,"syntology":null},{"paper":"/paper/coreference-aware-double-channel-attention","title":"Coreference-aware Double-channel Attention Network for Multi-party Dialogue Reading Comprehension","date":"2023-05-15","arxiv_id":"2305.08348","n_code_links":1,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/language-modelling","name":"Language Modelling","papers":31},{"task":"/task/language-modeling","name":"Language Modeling","papers":27},{"task":"/task/question-answering","name":"Question Answering","papers":17},{"task":"/task/sentence","name":"Sentence","papers":16},{"task":"/task/reading-comprehension","name":"Reading Comprehension","papers":10},{"task":"/task/masked-language-modeling","name":"Masked Language Modeling","papers":8},{"task":"/task/named-entity-recognition-1","name":"Named Entity Recognition","papers":7},{"task":"/task/named-entity-recognition-ner","name":"Named Entity Recognition (NER)","papers":7},{"task":"/task/natural-language-inference","name":"Natural Language Inference","papers":7},{"task":"/task/named-entity-recognition","name":"named-entity-recognition","papers":7},{"task":"/task/cg","name":"NER","papers":6},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":6},{"task":"/task/text-classification","name":"Text Classification","papers":6},{"task":"/task/classification-1","name":"Classification","papers":5},{"task":"/task/machine-reading-comprehension","name":"Machine Reading Comprehension","papers":5},{"task":"/task/multi-task-learning","name":"Multi-Task Learning","papers":5},{"task":"/task/task-2","name":"Task 2","papers":5},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":5},{"task":"/task/text-classification-1","name":"text-classification","papers":5},{"task":"/task/binary-classification","name":"Binary Classification","papers":4}],"tasks_shown":20,"n_tasks":136,"usage_by_year":[{"year":"2020","papers":19},{"year":"2021","papers":35},{"year":"2022","papers":29},{"year":"2023","papers":18},{"year":"2024","papers":11},{"year":"2025","papers":6}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/electra"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}