{"url":"/method/bert","slug":"bert","name":"BERT","full_name":"BERT","full_name_withheld":false,"description_markdown":"**BERT**, or Bidirectional Encoder Representations from Transformers, improves upon standard [Transformers](http://paperswithcode.com/method/transformer) by removing the unidirectionality constraint by using a *masked language model* (MLM) pre-training objective. The masked language model randomly masks some of the tokens from the input, and the objective is to predict the original vocabulary id of the masked word based only on its context. Unlike left-to-right language model pre-training, the MLM objective enables the representation to fuse the left and the right context, which allows us to pre-train a deep bidirectional Transformer. In addition to the masked language model, BERT uses a *next sentence prediction* task that jointly pre-trains text-pair representations. \r\n\r\nThere are two steps in BERT: *pre-training* and *fine-tuning*. During pre-training, the model is trained on unlabeled data over different pre-training tasks. For fine-tuning, the BERT model is first initialized with the pre-trained parameters, and all of the parameters are fine-tuned using labeled data from the downstream tasks. Each downstream task has separate fine-tuned models, even though they\r\nare initialized with the same pre-trained parameters.","description_state":"present","introduced_year":null,"introduced_by":{"title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","paper":"/paper/bert-pre-training-of-deep-bidirectional","first_author":"Jacob Devlin","n_authors":4,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/bert-pre-training-of-deep-bidirectional"},"source":{"url":"https://arxiv.org/abs/1810.04805v2","title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/google-research/bert","code_snippet_url_on_a_code_host":true,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Autoencoding Transformers","url":"/methods/category/autoencoding-transformers","pwc_aliases":[]},{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Transformers","url":"/methods/category/transformers","pwc_aliases":[]},{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Language Models","url":"/methods/category/language-models","pwc_aliases":[]}],"n_papers_tagged":6938,"archive_num_papers":6938,"papers_newest_first":[{"paper":"/paper/developing-visual-augmented-q-a-system-using","title":"Developing Visual Augmented Q&A System using Scalable Vision Embedding Retrieval & Late Interaction Re-ranker","date":"2025-07-16","arxiv_id":"2507.12378","n_code_links":1,"syntology":null},{"paper":"/paper/addressing-data-imbalance-in-transformer","title":"Addressing Data Imbalance in Transformer-Based Multi-Label Emotion Detection with Weighted Loss","date":"2025-07-15","arxiv_id":"2507.11384","n_code_links":1,"syntology":null},{"paper":null,"title":"Leveraging RAG-LLMs for Urban Mobility Simulation and Analysis","date":"2025-07-14","arxiv_id":"2507.10382","n_code_links":0,"syntology":null},{"paper":null,"title":"SentiDrop: A Multi Modal Machine Learning model for Predicting Dropout in Distance Learning","date":"2025-07-14","arxiv_id":"2507.10421","n_code_links":0,"syntology":null},{"paper":"/paper/orchestrator-agent-trust-a-modular-agentic-ai","title":"Orchestrator-Agent Trust: A Modular Agentic AI Visual Classification System with Trust-Aware Orchestration and RAG-Based Reasoning","date":"2025-07-09","arxiv_id":"2507.10571","n_code_links":1,"syntology":null},{"paper":null,"title":"The Dark Side of LLMs Agent-based Attacks for Complete Computer Takeover","date":"2025-07-09","arxiv_id":"2507.06850","n_code_links":0,"syntology":null},{"paper":null,"title":"SARA: Selective and Adaptive Retrieval-augmented Generation with Context Compression","date":"2025-07-08","arxiv_id":"2507.05633","n_code_links":0,"syntology":null},{"paper":null,"title":"AI Generated Text Detection Using Instruction Fine-tuned Large Language and Transformer-Based Models","date":"2025-07-07","arxiv_id":"2507.05157","n_code_links":0,"syntology":null},{"paper":null,"title":"CyberRAG: An agentic RAG cyber attack classification and reporting tool","date":"2025-07-03","arxiv_id":"2507.02424","n_code_links":0,"syntology":null},{"paper":null,"title":"Knowledge Protocol Engineering: A New Paradigm for AI in Domain-Specific Knowledge Work","date":"2025-07-03","arxiv_id":"2507.02760","n_code_links":0,"syntology":null},{"paper":"/paper/robustness-of-misinformation-classification","title":"Robustness of Misinformation Classification Systems to Adversarial Examples Through BeamAttack","date":"2025-06-30","arxiv_id":"2506.23661","n_code_links":1,"syntology":null},{"paper":null,"title":"Knowledge Augmented Finetuning Matters in both RAG and Agent Based Dialog Systems","date":"2025-06-28","arxiv_id":"2506.22852","n_code_links":0,"syntology":null},{"paper":null,"title":"ARAG: Agentic Retrieval Augmented Generation for Personalized Recommendation","date":"2025-06-27","arxiv_id":"2506.21931","n_code_links":0,"syntology":null},{"paper":"/paper/erarag-efficient-and-incremental-retrieval","title":"EraRAG: Efficient and Incremental Retrieval Augmented Generation for Growing Corpora","date":"2025-06-26","arxiv_id":"2506.20963","n_code_links":1,"syntology":null},{"paper":null,"title":"Leveraging LLM-Assisted Query Understanding for Live Retrieval-Augmented Generation","date":"2025-06-26","arxiv_id":"2506.21384","n_code_links":0,"syntology":null},{"paper":"/paper/psylite-technical-report","title":"PsyLite Technical Report","date":"2025-06-26","arxiv_id":"2506.21536","n_code_links":1,"syntology":null},{"paper":"/paper/response-quality-assessment-for-retrieval","title":"Response Quality Assessment for Retrieval-Augmented Generation via Conditional Conformal Factuality","date":"2025-06-26","arxiv_id":"2506.20978","n_code_links":1,"syntology":{"ran":0,"of":3,"unverified":3,"pointer_only":3}},{"paper":null,"title":"AI Assistants to Enhance and Exploit the PETSc Knowledge Base","date":"2025-06-25","arxiv_id":"2506.20608","n_code_links":0,"syntology":null},{"paper":null,"title":"CCRS: A Zero-Shot LLM-as-a-Judge Framework for Comprehensive RAG Evaluation","date":"2025-06-25","arxiv_id":"2506.20128","n_code_links":0,"syntology":null},{"paper":null,"title":"Engineering RAG Systems for Real-World Applications: Design, Development, and Evaluation","date":"2025-06-25","arxiv_id":"2506.20869","n_code_links":0,"syntology":null},{"paper":null,"title":"Knowledge-Aware Diverse Reranking for Cross-Source Question Answering","date":"2025-06-25","arxiv_id":"2506.20476","n_code_links":0,"syntology":null},{"paper":null,"title":"Memento: Note-Taking for Your Future Self","date":"2025-06-25","arxiv_id":"2506.20642","n_code_links":0,"syntology":null},{"paper":null,"title":"Accurate and Energy Efficient: Local Retrieval-Augmented Generation Models Outperform Commercial Large Language Models in Medical Tasks","date":"2025-06-24","arxiv_id":"2506.20009","n_code_links":0,"syntology":null},{"paper":null,"title":"Controlled Retrieval-augmented Context Evaluation for Long-form RAG","date":"2025-06-24","arxiv_id":"2506.20051","n_code_links":0,"syntology":null},{"paper":null,"title":"Inference Scaled GraphRAG: Improving Multi Hop Question Answering on Knowledge Graphs","date":"2025-06-24","arxiv_id":"2506.19967","n_code_links":0,"syntology":null},{"paper":null,"title":"KunLunBaizeRAG: Reinforcement Learning Driven Inference Performance Leap for Large Language Models","date":"2025-06-24","arxiv_id":"2506.19466","n_code_links":0,"syntology":null},{"paper":null,"title":"Unlocking Insights Addressing Alcohol Inference Mismatch through Database-Narrative Alignment","date":"2025-06-24","arxiv_id":"2506.19342","n_code_links":0,"syntology":null},{"paper":null,"title":"An Audio-centric Multi-task Learning Framework for Streaming Ads Targeting on Spotify","date":"2025-06-23","arxiv_id":"2506.18735","n_code_links":0,"syntology":null},{"paper":null,"title":"Semantic similarity estimation for domain specific data using BERT and other techniques","date":"2025-06-23","arxiv_id":"2506.18602","n_code_links":0,"syntology":null},{"paper":null,"title":"T-CPDL: A Temporal Causal Probabilistic Description Logic for Developing Logic-RAG Agent","date":"2025-06-23","arxiv_id":"2506.18559","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/rag","name":"RAG","papers":1288},{"task":"/task/retrieval","name":"Retrieval","papers":1260},{"task":"/task/language-modelling","name":"Language Modelling","papers":1119},{"task":"/task/retrieval-augmented-generation","name":"Retrieval-augmented Generation","papers":1092},{"task":"/task/language-modeling","name":"Language Modeling","papers":903},{"task":"/task/question-answering","name":"Question Answering","papers":751},{"task":"/task/sentence","name":"Sentence","papers":730},{"task":"/task/sentiment-analysis","name":"Sentiment Analysis","papers":406},{"task":"/task/text-classification","name":"Text Classification","papers":387},{"task":"/task/text-classification-1","name":"text-classification","papers":342},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":305},{"task":"/task/classification-1","name":"Classification","papers":261},{"task":"/task/word-embeddings","name":"Word Embeddings","papers":258},{"task":"/task/information-retrieval","name":"Information Retrieval","papers":243},{"task":"/task/named-entity-recognition-1","name":"Named Entity Recognition","papers":242},{"task":"/task/named-entity-recognition-ner","name":"Named Entity Recognition (NER)","papers":231},{"task":"/task/named-entity-recognition","name":"named-entity-recognition","papers":231},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":229},{"task":"/task/large-language-model","name":"Large Language Model","papers":221},{"task":"/task/articles","name":"Articles","papers":188}],"tasks_shown":20,"n_tasks":1187,"usage_by_year":[{"year":"2018","papers":6},{"year":"2019","papers":576},{"year":"2020","papers":1234},{"year":"2021","papers":1288},{"year":"2022","papers":810},{"year":"2023","papers":831},{"year":"2024","papers":1375},{"year":"2025","papers":818}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/bert"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}