{"url":"/method/squared-relu","slug":"squared-relu","name":"Squared ReLU","full_name":"Squared ReLU","full_name_withheld":false,"description_markdown":"**Squared ReLU** is an activation function used in the [Primer](https://paperswithcode.com/method/primer) architecture in the feedforward block of the [Transformer](https://paperswithcode.com/methods/category/transformers) layer. It is simply squared [ReLU](https://paperswithcode.com/method/relu) activations.\r\n\r\nThe effectiveness of higher order polynomials can also be observed in other effective [Transformer](https://paperswithcode.com/method/transformer) nonlinearities, such as [GLU](https://paperswithcode.com/method/glu) variants like [ReGLU](https://paperswithcode.com/method/reglu) and point-wise activations like [approximate GELU](https://paperswithcode.com/method/gelu). However, squared ReLU has drastically different asymptotics as $x \\rightarrow \\inf$ compared to the most commonly used activation functions: [ReLU](https://paperswithcode.com/method/relu), [GELU](https://paperswithcode.com/method/gelu) and [Swish](https://paperswithcode.com/method/swish). Squared ReLU does have significant overlap with ReGLU and in fact is equivalent when ReGLU’s $U$ and $V$ weight matrices are the same and squared ReLU is immediately preceded by a linear transformation with weight matrix $U$. This leads the authors to believe that squared ReLUs capture the benefits of these GLU variants, while being simpler, without additional parameters, and delivering better quality.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Primer: Searching for Efficient Transformers for Language Modeling","paper":"/paper/primer-searching-for-efficient-transformers","first_author":"David R. So","n_authors":6,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/primer-searching-for-efficient-transformers"},"source":{"url":"https://arxiv.org/abs/2109.08668v2","title":"Primer: Searching for Efficient Transformers for Language Modeling","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Activation Functions","url":"/methods/category/activation-functions","pwc_aliases":[]}],"n_papers_tagged":15,"archive_num_papers":15,"papers_newest_first":[{"paper":null,"title":"SAFEPATH: Preventing Harmful Reasoning in Chain-of-Thought via Early Alignment","date":"2025-05-20","arxiv_id":"2505.14667","n_code_links":0,"syntology":null},{"paper":null,"title":"A review of DNA restriction-free overlapping sequence cloning techniques for synthetic biology","date":"2025-05-06","arxiv_id":"2505.03681","n_code_links":0,"syntology":null},{"paper":null,"title":"Primer C-VAE: An interpretable deep learning primer design method to detect emerging virus variants","date":"2025-03-03","arxiv_id":"2503.01459","n_code_links":0,"syntology":null},{"paper":"/paper/deriving-activation-functions-via-integration","title":"Deriving Activation Functions Using Integration","date":"2024-11-20","arxiv_id":"2411.13010","n_code_links":1,"syntology":null},{"paper":"/paper/solving-oscillator-odes-via-soft-constrained","title":"Characteristic Performance Study on Solving Oscillator ODEs via Soft-constrained Physics-informed Neural Network with Small Data","date":"2024-08-19","arxiv_id":"2408.11077","n_code_links":1,"syntology":null},{"paper":null,"title":"The curious case of A31P, a topology-switching mutant of the Repressor of Primer protein : A molecular dynamics study of its folding and misfolding","date":"2024-04-01","arxiv_id":"2404.01405","n_code_links":0,"syntology":null},{"paper":null,"title":"Brainformers: Trading Simplicity for Efficiency","date":"2023-05-29","arxiv_id":"2306.00008","n_code_links":0,"syntology":null},{"paper":"/paper/the-effects-of-political-martyrdom-on","title":"The Effects of Political Martyrdom on Election Results: The Assassination of Abe","date":"2023-05-29","arxiv_id":"2305.18004","n_code_links":1,"syntology":null},{"paper":null,"title":"Towards NeuroAI: Introducing Neuronal Diversity into Artificial Neural Networks","date":"2023-01-23","arxiv_id":"2301.09245","n_code_links":0,"syntology":null},{"paper":"/paper/n-grammer-augmenting-transformers-with-latent-1","title":"N-Grammer: Augmenting Transformers with latent n-grams","date":"2022-07-13","arxiv_id":"2207.06366","n_code_links":2,"syntology":{"ran":0,"of":6,"unverified":6,"pointer_only":0}},{"paper":null,"title":"Piecewise Linear Neural Networks and Deep Learning","date":"2022-06-18","arxiv_id":"2206.09149","n_code_links":0,"syntology":null},{"paper":null,"title":"Enriching and Characterizing T-Cell Repertoires from 3' Barcoded Single-Cell Whole Transcriptome Amplification Products","date":"2022-03-21","arxiv_id":"2203.11266","n_code_links":0,"syntology":null},{"paper":null,"title":"Searching for Efficient Transformers for Language Modeling","date":"2021-12-01","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":null,"title":"N-grammer: Augmenting Transformers with latent n-grams","date":"2021-11-16","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":"/paper/primer-searching-for-efficient-transformers","title":"Primer: Searching for Efficient Transformers for Language Modeling","date":"2021-09-17","arxiv_id":"2109.08668","n_code_links":4,"syntology":{"ran":3,"of":3,"unverified":0,"pointer_only":3}}],"papers_shown":15,"tasks":[{"task":"/task/language-modeling","name":"Language Modeling","papers":4},{"task":"/task/language-modelling","name":"Language Modelling","papers":4},{"task":null,"name":"CPU","papers":1},{"task":"/task/common-sense-reasoning","name":"Common Sense Reasoning","papers":1},{"task":"/task/coreference-resolution","name":"Coreference Resolution","papers":1},{"task":"/task/deep-learning","name":"Deep Learning","papers":1},{"task":"/task/diversity","name":"Diversity","papers":1},{"task":"/task/epidemiology","name":"Epidemiology","papers":1},{"task":null,"name":"GPU","papers":1},{"task":"/task/natural-language-inference","name":"Natural Language Inference","papers":1},{"task":"/task/protein-structure-prediction","name":"Protein Structure Prediction","papers":1},{"task":"/task/question-answering","name":"Question Answering","papers":1},{"task":"/task/safety-alignment","name":"Safety Alignment","papers":1},{"task":"/task/sentiment-analysis","name":"Sentiment Analysis","papers":1},{"task":"/task/specificity","name":"Specificity","papers":1},{"task":"/task/tar","name":"TAR","papers":1},{"task":"/task/text-classification","name":"Text Classification","papers":1},{"task":"/task/word-sense-disambiguation","name":"Word Sense Disambiguation","papers":1}],"tasks_shown":18,"n_tasks":18,"usage_by_year":[{"year":"2021","papers":3},{"year":"2022","papers":3},{"year":"2023","papers":3},{"year":"2024","papers":3},{"year":"2025","papers":3}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/squared-relu"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}