Methods › General › Activation Functions › Squared ReLU
Squared ReLU
Introduced by David R. So et al. in Primer: Searching for Efficient Transformers for Language Modeling
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
Squared ReLU is an activation function used in the Primer architecture in the feedforward block of the Transformer layer. It is simply squared ReLU activations.
The effectiveness of higher order polynomials can also be observed in other effective Transformer nonlinearities, such as GLU variants like ReGLU and point-wise activations like approximate GELU. However, squared ReLU has drastically different asymptotics as x →inf compared to the most commonly used activation functions: ReLU, GELU and Swish. Squared ReLU does have significant overlap with ReGLU and in fact is equivalent when ReGLU’s U and V weight matrices are the same and squared ReLU is immediately preceded by a linear transformation with weight matrix U. This leads the authors to believe that squared ReLUs capture the benefits of these GLU variants, while being simpler, without additional parameters, and delivering better quality.
Papers archive 2025-07-28
15 shown of 15, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment 20 May 2025 · 0 repositories · arXiv:2505.14667
-
A review of DNA restriction-free overlapping sequence cloning techniques for synthetic biology 6 May 2025 · 0 repositories · arXiv:2505.03681
-
Primer C-VAE: An interpretable deep learning primer design method to detect emerging virus variants 3 Mar 2025 · 0 repositories · arXiv:2503.01459
-
Deriving Activation Functions Using Integration 20 Nov 2024 · 1 repository · arXiv:2411.13010
-
Characteristic Performance Study on Solving Oscillator ODEs via Soft-constrained Physics-informed Neural Network with Small Data 19 Aug 2024 · 1 repository · arXiv:2408.11077
-
The curious case of A31P, a topology-switching mutant of the Repressor of Primer protein : A molecular dynamics study of its folding and misfolding 1 Apr 2024 · 0 repositories · arXiv:2404.01405
-
Brainformers: Trading Simplicity for Efficiency 29 May 2023 · 0 repositories · arXiv:2306.00008
-
The Effects of Political Martyrdom on Election Results: The Assassination of Abe 29 May 2023 · 1 repository · arXiv:2305.18004
-
Towards NeuroAI: Introducing Neuronal Diversity into Artificial Neural Networks 23 Jan 2023 · 0 repositories · arXiv:2301.09245
-
N-Grammer: Augmenting Transformers with latent n-grams 13 Jul 2022 · 2 repositories · arXiv:2207.06366Syntology ran 0 of 6 samples · 6 unverified
-
Piecewise Linear Neural Networks and Deep Learning 18 Jun 2022 · 0 repositories · arXiv:2206.09149
-
Enriching and Characterizing T-Cell Repertoires from 3' Barcoded Single-Cell Whole Transcriptome Amplification Products 21 Mar 2022 · 0 repositories · arXiv:2203.11266
-
Searching for Efficient Transformers for Language Modeling 1 Dec 2021 · 0 repositories
-
N-grammer: Augmenting Transformers with latent n-grams 16 Nov 2021 · 0 repositories
-
Primer: Searching for Efficient Transformers for Language Modeling 17 Sep 2021 · 4 repositories · arXiv:2109.08668Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)
Tasks archive 2025-07-28
18 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections