Methods › General › Attention › Multi-Query Attention
Multi-Query Attention
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Multi-head attention consists of multiple attention layers (heads) in parallel with different linear transformations on the queries, keys, values and outputs. Multi-query attention is identical except that the different heads share a single set of keys and values.
Papers archive 2025-07-28
13 shown of 13, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
The Nature of Mathematical Modeling and Probabilistic Optimization Engineering in Generative AI 24 Oct 2024 · 0 repositories · arXiv:2410.18441
-
Weighted Grouped Query Attention in Transformers 15 Jul 2024 · 0 repositories · arXiv:2407.10855
-
MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding 13 Jun 2024 · 1 repository · arXiv:2406.09297
-
Effectively Compress KV Heads for LLM 11 Jun 2024 · 0 repositories · arXiv:2406.07056
-
QCQA: Quality and Capacity-aware grouped Query Attention 8 Jun 2024 · 0 repositories · arXiv:2406.10247
-
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention 21 May 2024 · 2 repositories · arXiv:2405.12981
-
Bifurcated Attention: Accelerating Massively Parallel Decoding with Shared Prefixes in LLMs 13 Mar 2024 · 0 repositories · arXiv:2403.08845
-
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models 29 Feb 2024 · 4 repositories · arXiv:2402.19427Syntology ran 0 of 10 samples · 10 unverified
-
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints 22 May 2023 · 4 repositories · arXiv:2305.13245Syntology ran 4 of 5 samples · 1 unverified · 3 pointer-only (licence)
-
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness 27 May 2022 · 13 repositories · arXiv:2205.14135Syntology ran 9 of 30 samples · 21 unverified · 1 pointer-only (licence)
-
PaLM: Scaling Language Modeling with Pathways 5 Apr 2022 · 7 repositories · arXiv:2204.02311Syntology ran 30 of 37 samples · 7 unverified
-
Fast Transformer Decoding: One Write-Head is All You Need 6 Nov 2019 · 4 repositories · arXiv:1911.02150Syntology ran 0 of 3 samples · 3 unverified
-
Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm 5 Dec 2017 · 62 repositories · arXiv:1712.01815Syntology ran 13 of 17 samples · 4 unverified · 9 pointer-only (licence)
Tasks archive 2025-07-28
20 shown of 44 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections