Methods › Natural Language Processing › Synthesized Attention Mechanisms › Factorized Random Synthesized Attention

Factorized Random Synthesized Attention

1 paper tagged archive 2025-07-28

Introduced by Yi Tay et al. in Synthesizer: Rethinking Self-Attention in Transformer Models

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

Factorized Random Synthesized Attention, introduced with the Synthesizer architecture, is similar to factorized dense synthesized attention but for random synthesizers. Letting R being a randomly initialized matrix, we factorize R into low rank matrices R₁, R₂ ∈ℝ^(l xk) in the attention function:

Y = Softmax(R₁R₂ᵀ)G(X) .

Here G(.) is a parameterized function that is equivalent to V in Scaled Dot-Product Attention.

For each head, the factorization reduces the parameter costs from l² to 2(lk) where k << l and hence helps prevent overfitting. In practice, we use a small value of k = 8.

The basic idea of a Random Synthesizer is to not rely on pairwise token interactions or any information from individual token but rather to learn a task-specific alignment that works well globally across many samples.

PaperSource

Papers archive 2025-07-28

1 shown of 1, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

10 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Abstractive Text Summarization1
Dialogue Generation1
Document Summarization1
Language Modeling1
Language Modelling1
Linguistic Acceptability1
Machine Translation1
Semantic Textual Similarity1
Text Generation1
Translation1

Usage over time archive 2025-07-28

Papers per year tagged with Factorized Random Synthesized Attention: 2020 to 2020, peak 1 1 0 2020: 1 paper 2020
Papers per year the archive tags with this method, by the paper's archive date (1 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Synthesized Attention MechanismsAttention Mechanisms

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections