Methods › General › Attention › Grouped-query attention

Grouped-query attention

19 papers tagged archive 2025-07-28

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

Grouped-query attention an interpolation of multi-query and multi-head attention that achieves quality close to multi-head at comparable speed to multi-query attention.

Source: GQA: Training Generalized Multi-Query Transformer Models...

Papers archive 2025-07-28

19 shown of 19, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 35 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Language Modelling6
Language Modeling5
Large Language Model3
Question Answering3
4k2
Arithmetic Reasoning2
Chatbot2
Code Generation2
Decoder2
GPU2
Mamba2
Math Word Problem Solving2
Multi-task Language Understanding2
Sentence Completion2
16k1
8k1
Common Sense Reasoning1
Computational Efficiency1
Document Classification1
HumanEval1

Usage over time archive 2025-07-28

Papers per year tagged with Grouped-query attention: 2022 to 2025, peak 10 10 0 2022: 1 paper 2022 2023: 3 papers 2023 2024: 10 papers 2024 2025: 5 papers 2025
Papers per year the archive tags with this method, by the paper's archive date (19 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Attention

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections