Methods › General › Attention › Grouped-query attention
Grouped-query attention
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Grouped-query attention an interpolation of multi-query and multi-head attention that achieves quality close to multi-head at comparable speed to multi-query attention.
Papers archive 2025-07-28
19 shown of 19, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Hardware-Efficient Attention for Fast Decoding 27 May 2025 · 2 repositories · arXiv:2505.21487
-
Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought 21 May 2025 · 0 repositories · arXiv:2505.15431
-
Towards Economical Inference: Enabling DeepSeek's Multi-Head Latent Attention in Any Transformer-based LLMs 20 Feb 2025 · 1 repository · arXiv:2502.14837Syntology ran 0 of 15 samples · 15 unverified
-
FastKV: KV Cache Compression for Fast Long-Context Processing with Token-Selective Propagation 3 Feb 2025 · 1 repository · arXiv:2502.01068
-
Sigma: Differential Rescaling of Query, Key and Value for Efficient Language Models 23 Jan 2025 · 0 repositories · arXiv:2501.13629
-
UNComp: Uncertainty-Aware Long-Context Compressor for Efficient Large Language Model Inference 4 Oct 2024 · 0 repositories · arXiv:2410.03090
-
Weighted Grouped Query Attention in Transformers 15 Jul 2024 · 0 repositories · arXiv:2407.10855
-
MLKV: Multi-Layer Key-Value Heads for Memory Efficient Transformer Decoding 13 Jun 2024 · 1 repository · arXiv:2406.09297
-
Effectively Compress KV Heads for LLM 11 Jun 2024 · 0 repositories · arXiv:2406.07056
-
Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling 11 Jun 2024 · 2 repositories · arXiv:2406.07522Syntology ran 1 of 1 samples · 0 unverified
-
Reducing Transformer Key-Value Cache Size with Cross-Layer Attention 21 May 2024 · 2 repositories · arXiv:2405.12981
-
CodeShell Technical Report 23 Mar 2024 · 0 repositories · arXiv:2403.15747
-
Vi-Mistral-X: Building a Vietnamese Language Model with Advanced Continual Pre-training 20 Mar 2024 · 0 repositories · arXiv:2403.15470
-
Dynamic Memory Compression: Retrofitting LLMs for Accelerated Inference 14 Mar 2024 · 1 repository · arXiv:2403.09636
-
MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases 22 Feb 2024 · 4 repositories · arXiv:2402.14905Syntology ran 11 of 28 samples · 17 unverified · 20 pointer-only (licence)
-
Mistral 7B 10 Oct 2023 · 6 repositories · arXiv:2310.06825Syntology ran 9 of 11 samples · 2 unverified · 1 pointer-only (licence)
-
Llama 2: Open Foundation and Fine-Tuned Chat Models 18 Jul 2023 · 19 repositories · arXiv:2307.09288Syntology ran 31 of 52 samples · 21 unverified · 16 pointer-only (licence)
-
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints 22 May 2023 · 4 repositories · arXiv:2305.13245Syntology ran 4 of 5 samples · 1 unverified · 3 pointer-only (licence)
-
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness 27 May 2022 · 13 repositories · arXiv:2205.14135Syntology ran 9 of 30 samples · 21 unverified · 1 pointer-only (licence)
Tasks archive 2025-07-28
20 shown of 35 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections