Methods › General › Attention › Multi-Query Attention

Multi-Query Attention

13 papers tagged archive 2025-07-28

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

Multi-head attention consists of multiple attention layers (heads) in parallel with different linear transformations on the queries, keys, values and outputs. Multi-query attention is identical except that the different heads share a single set of keys and values.

Source: Fast Transformer Decoding: One Write-Head is All You Need

Papers archive 2025-07-28

13 shown of 13, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 44 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Language Modelling6
Language Modeling3
Decoder2
16k1
4k1
8k1
All1
Answer Generation1
Auto Debugging1
Code Generation1
Common Sense Reasoning1
Coreference Resolution1
Cross-Lingual Question Answering1
Document Classification1
Few-Shot Learning1
GPU1
Game of Chess1
Game of Go1
Game of Shogi1
General Reinforcement Learning1

Usage over time archive 2025-07-28

Papers per year tagged with Multi-Query Attention: 2017 to 2024, peak 8 8 0 2017: 1 paper 2017 2018: 0 papers 2018 2019: 1 paper 2019 2020: 0 papers 2020 2021: 0 papers 2021 2022: 2 papers 2022 2023: 1 paper 2023 2024: 8 papers 2024
Papers per year the archive tags with this method, by the paper's archive date (13 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Attention

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections