Methods › General › Attention › Attention Sinks
Attention Sinks
Introduced by Guangxuan Xiao et al. in Efficient Streaming Language Models with Attention Sinks
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
The archive carries only a placeholder description for this method.
Papers archive 2025-07-28
13 shown of 13, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
MARS-Bench: A Multi-turn Athletic Real-world Scenario Benchmark for Dialogue Evaluation 27 May 2025 · 0 repositories · arXiv:2505.23810
-
Analysis of Attention in Video Diffusion Transformers 14 Apr 2025 · 0 repositories · arXiv:2504.10317
-
Why do LLMs attend to the first token? 3 Apr 2025 · 1 repository · arXiv:2504.02732Syntology ran 0 of 7 samples · 7 unverified
-
Interpreting the Repeated Token Phenomenon in Large Language Models 11 Mar 2025 · 1 repository · arXiv:2503.08908
-
Attention Sinks and Outlier Features: A 'Catch, Tag, and Release' Mechanism for Embeddings 2 Feb 2025 · 0 repositories · arXiv:2502.00919
-
Task-KV: Task-aware KV Cache Optimization via Semantic Differentiation of Attention Heads 25 Jan 2025 · 0 repositories · arXiv:2501.15113
-
Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models 21 Dec 2024 · 0 repositories · arXiv:2412.16545
-
Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs 15 Nov 2024 · 0 repositories · arXiv:2411.09968
-
Value Residual Learning For Alleviating Attention Concentration In Transformers 23 Oct 2024 · 1 repository · arXiv:2410.17897Syntology ran 3 of 12 samples · 9 unverified
-
When Attention Sink Emerges in Language Models: An Empirical View 14 Oct 2024 · 1 repository · arXiv:2410.10781Syntology ran 3 of 3 samples · 0 unverified
-
Does RoBERTa Perform Better than BERT in Continual Learning: An Attention Sink Perspective 8 Oct 2024 · 1 repository · arXiv:2410.05648
-
Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration 22 Jun 2024 · 1 repository · arXiv:2406.15765Syntology ran 5 of 13 samples · 8 unverified · 13 pointer-only (licence)
-
Efficient Streaming Language Models with Attention Sinks 29 Sep 2023 · 6 repositories · arXiv:2309.17453Syntology ran 8 of 11 samples · 3 unverified · 2 pointer-only (licence)
Tasks archive 2025-07-28
11 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
| Task | Papers |
|---|---|
| Continual Learning | 1 |
| Decoder | 1 |
| Dialogue Evaluation | 1 |
| Diversity | 1 |
| Hallucination | 1 |
| Language Modeling | 1 |
| Language Modelling | 1 |
| Model Compression | 1 |
| Quantization | 1 |
| TAG | 1 |
| Video Editing | 1 |
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections