{"url":"/method/convbert","slug":"convbert","name":"ConvBERT","full_name":"ConvBERT","full_name_withheld":false,"description_markdown":"**ConvBERT** is a modification on the [BERT](https://paperswithcode.com/method/bert) architecture which uses a [span-based dynamic convolution](https://paperswithcode.com/method/span-based-dynamic-convolution) to replace self-attention heads to directly model local dependencies. Specifically a new [mixed attention module](https://paperswithcode.com/method/mixed-attention-block) replaces the [self-attention modules](https://paperswithcode.com/method/scaled) in BERT, which leverages the advantages of [convolution](https://paperswithcode.com/method/convolution) to better capture local dependency. Additionally, a new span-based dynamic convolution operation is used to utilize multiple input tokens to dynamically generate the convolution kernel. Lastly, ConvBERT also incorporates some new model designs including the bottleneck attention and grouped linear operator for the feed-forward module (reducing the number of parameters).","description_state":"present","introduced_year":null,"introduced_by":{"title":"ConvBERT: Improving BERT with Span-based Dynamic Convolution","paper":"/paper/convbert-improving-bert-with-span-based","first_author":"Zi-Hang Jiang","n_authors":6,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/convbert-improving-bert-with-span-based"},"source":{"url":"https://arxiv.org/abs/2008.02496v3","title":"ConvBERT: Improving BERT with Span-based Dynamic Convolution","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Autoencoding Transformers","url":"/methods/category/autoencoding-transformers","pwc_aliases":[]},{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Transformers","url":"/methods/category/transformers","pwc_aliases":[]}],"n_papers_tagged":5,"archive_num_papers":5,"papers_newest_first":[{"paper":null,"title":"Beyond Simple Concatenation: Fairly Assessing PLM Architectures for Multi-Chain Protein-Protein Interactions Prediction","date":"2025-05-26","arxiv_id":"2505.20036","n_code_links":0,"syntology":null},{"paper":"/paper/navigating-nuance-in-quest-for-political","title":"Navigating Nuance: In Quest for Political Truth","date":"2025-01-01","arxiv_id":"2501.00782","n_code_links":1,"syntology":null},{"paper":null,"title":"ChatGPT v.s. Media Bias: A Comparative Study of GPT-3.5 and Fine-tuned Language Models","date":"2024-03-29","arxiv_id":"2403.20158","n_code_links":0,"syntology":null},{"paper":"/paper/transformer-based-punctuation-restoration-for","title":"Transformer Based Punctuation Restoration for Turkish","date":"2023-09-15","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/convbert-improving-bert-with-span-based","title":"ConvBERT: Improving BERT with Span-based Dynamic Convolution","date":"2020-08-06","arxiv_id":"2008.02496","n_code_links":8,"syntology":null}],"papers_shown":5,"tasks":[{"task":"/task/automatic-speech-recognition-2","name":"Automatic Speech Recognition","papers":1},{"task":"/task/automatic-speech-recognition","name":"Automatic Speech Recognition (ASR)","papers":1},{"task":"/task/bias-detection","name":"Bias Detection","papers":1},{"task":"/task/drug-discovery","name":"Drug Discovery","papers":1},{"task":"/task/language-modeling","name":"Language Modeling","papers":1},{"task":"/task/language-modelling","name":"Language Modelling","papers":1},{"task":"/task/misinformation","name":"Misinformation","papers":1},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":1},{"task":"/task/punctuation-restoration","name":"Punctuation Restoration","papers":1},{"task":"/task/speech-recognition","name":"Speech Recognition","papers":1},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":1},{"task":"/task/speech-recognition-1","name":"speech-recognition","papers":1}],"tasks_shown":12,"n_tasks":12,"usage_by_year":[{"year":"2020","papers":1},{"year":"2023","papers":1},{"year":"2024","papers":1},{"year":"2025","papers":2}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/convbert"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}