{"url":"/method/mixed-attention-block","slug":"mixed-attention-block","name":"Mixed Attention Block","full_name":"Mixed Attention Block","full_name_withheld":false,"description_markdown":"**Mixed Attention Block** is an attention module used in the [ConvBERT](https://paperswithcode.com/method/convbert) architecture. It is a mixture of [self-attention](https://paperswithcode.com/method/scaled) and [span-based dynamic convolution](https://paperswithcode.com/method/span-based-dynamic-convolution) (highlighted in pink). They share the same Query but use different Key to generate the attention map and [convolution](https://paperswithcode.com/method/convolution) kernel respectively. The number of attention heads is reducing by directly projecting the input to a smaller embedding space to form a bottleneck structure for self-attention and span-based dynamic convolution. Dimensions of the input and output of some blocks are labeled on the left top corner to illustrate the overall framework, where $d$ is the embedding size of the input and $\\gamma$ is the reduction ratio.","description_state":"present","introduced_year":null,"introduced_by":{"title":"ConvBERT: Improving BERT with Span-based Dynamic Convolution","paper":"/paper/convbert-improving-bert-with-span-based","first_author":"Zi-Hang Jiang","n_authors":6,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/convbert-improving-bert-with-span-based"},"source":{"url":"https://arxiv.org/abs/2008.02496v3","title":"ConvBERT: Improving BERT with Span-based Dynamic Convolution","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Attention Modules","url":"/methods/category/attention-modules","pwc_aliases":[]}],"n_papers_tagged":6,"archive_num_papers":6,"papers_newest_first":[{"paper":null,"title":"Beyond Simple Concatenation: Fairly Assessing PLM Architectures for Multi-Chain Protein-Protein Interactions Prediction","date":"2025-05-26","arxiv_id":"2505.20036","n_code_links":0,"syntology":null},{"paper":"/paper/navigating-nuance-in-quest-for-political","title":"Navigating Nuance: In Quest for Political Truth","date":"2025-01-01","arxiv_id":"2501.00782","n_code_links":1,"syntology":null},{"paper":null,"title":"ChatGPT v.s. Media Bias: A Comparative Study of GPT-3.5 and Fine-tuned Language Models","date":"2024-03-29","arxiv_id":"2403.20158","n_code_links":0,"syntology":null},{"paper":"/paper/m-3-net-multilevel-mixed-and-multistage","title":"M$^3$Net: Multilevel, Mixed and Multistage Attention Network for Salient Object Detection","date":"2023-09-15","arxiv_id":"2309.08365","n_code_links":1,"syntology":null},{"paper":"/paper/transformer-based-punctuation-restoration-for","title":"Transformer Based Punctuation Restoration for Turkish","date":"2023-09-15","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/convbert-improving-bert-with-span-based","title":"ConvBERT: Improving BERT with Span-based Dynamic Convolution","date":"2020-08-06","arxiv_id":"2008.02496","n_code_links":8,"syntology":null}],"papers_shown":6,"tasks":[{"task":"/task/automatic-speech-recognition-2","name":"Automatic Speech Recognition","papers":1},{"task":"/task/automatic-speech-recognition","name":"Automatic Speech Recognition (ASR)","papers":1},{"task":"/task/bias-detection","name":"Bias Detection","papers":1},{"task":"/task/drug-discovery","name":"Drug Discovery","papers":1},{"task":"/task/language-modeling","name":"Language Modeling","papers":1},{"task":"/task/language-modelling","name":"Language Modelling","papers":1},{"task":"/task/misinformation","name":"Misinformation","papers":1},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":1},{"task":"/task/object-detection","name":"Object Detection","papers":1},{"task":"/task/punctuation-restoration","name":"Punctuation Restoration","papers":1},{"task":"/task/salient-object-detection","name":"RGB Salient Object Detection","papers":1},{"task":"/task/salient-object-detection-1","name":"Salient Object Detection","papers":1},{"task":"/task/speech-recognition","name":"Speech Recognition","papers":1},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":1},{"task":"/task/object-detection-1","name":"object-detection","papers":1},{"task":"/task/speech-recognition-1","name":"speech-recognition","papers":1}],"tasks_shown":16,"n_tasks":16,"usage_by_year":[{"year":"2020","papers":1},{"year":"2023","papers":2},{"year":"2024","papers":1},{"year":"2025","papers":2}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/mixed-attention-block"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}