{"url":"/method/silu","slug":"silu","name":"SiLU","full_name":"Sigmoid Linear Unit","full_name_withheld":false,"description_markdown":"** Sigmoid Linear Units**, or **SiLUs**, are activation functions for\r\nneural networks. The activation of the SiLU is computed by the sigmoid function multiplied by its input, or $$ x\\sigma(x).$$\r\n\r\nSee [Gaussian Error Linear Units](https://arxiv.org/abs/1606.08415) ([GELUs](https://paperswithcode.com/method/gelu)) where the SiLU was originally coined, and see [Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning](https://arxiv.org/abs/1702.03118) and [Swish: a Self-Gated Activation Function](https://arxiv.org/abs/1710.05941v1) where the SiLU was experimented with later.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning","paper":"/paper/sigmoid-weighted-linear-units-for-neural","first_author":"Stefan Elfwing","n_authors":3,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/sigmoid-weighted-linear-units-for-neural"},"source":{"url":"http://arxiv.org/abs/1702.03118v3","title":"Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/hendrycks/GELUs/blob/master/twitter_pos.py#L178","code_snippet_url_on_a_code_host":true,"categories":[{"area":"General","area_id":"general","collection":"Activation Functions","url":"/methods/category/activation-functions","pwc_aliases":[]}],"n_papers_tagged":30,"archive_num_papers":30,"papers_newest_first":[{"paper":"/paper/mvnet-hyperspectral-remote-sensing-image","title":"MVNet: Hyperspectral Remote Sensing Image Classification Based on Hybrid Mamba-Transformer Vision Backbone Architecture","date":"2025-07-06","arxiv_id":"2507.04409","n_code_links":1,"syntology":null},{"paper":"/paper/deriving-activation-functions-via-integration","title":"Deriving Activation Functions Using Integration","date":"2024-11-20","arxiv_id":"2411.13010","n_code_links":1,"syntology":null},{"paper":"/paper/sparsing-law-towards-large-language-models","title":"Sparsing Law: Towards Large Language Models with Greater Activation Sparsity","date":"2024-11-04","arxiv_id":"2411.02335","n_code_links":1,"syntology":{"ran":1,"of":5,"unverified":4,"pointer_only":0}},{"paper":null,"title":"ActNAS : Generating Efficient YOLO Models using Activation NAS","date":"2024-10-11","arxiv_id":"2410.10887","n_code_links":0,"syntology":null},{"paper":"/paper/unsegarmanet-unsupervised-image-segmentation","title":"UnSeGArmaNet: Unsupervised Image Segmentation using Graph Neural Networks with Convolutional ARMA Filters","date":"2024-10-08","arxiv_id":"2410.06114","n_code_links":1,"syntology":null},{"paper":null,"title":"BrainTransformers: SNN-LLM","date":"2024-10-03","arxiv_id":"2410.14687","n_code_links":0,"syntology":null},{"paper":null,"title":"Efficient Privacy-Preserving KAN Inference Using Homomorphic Encryption","date":"2024-09-12","arxiv_id":"2409.07751","n_code_links":0,"syntology":null},{"paper":null,"title":"CipherDM: Secure Three-Party Inference for Diffusion Model Sampling","date":"2024-09-09","arxiv_id":"2409.05414","n_code_links":0,"syntology":null},{"paper":null,"title":"On Expressive Power of Quantized Neural Networks under Fixed-Point Arithmetic","date":"2024-08-30","arxiv_id":"2409.00297","n_code_links":0,"syntology":null},{"paper":"/paper/reducing-fine-tuning-memory-overhead-by","title":"Reducing Fine-Tuning Memory Overhead by Approximate and Memory-Sharing Backpropagation","date":"2024-06-24","arxiv_id":"2406.16282","n_code_links":1,"syntology":{"ran":10,"of":21,"unverified":11,"pointer_only":0}},{"paper":null,"title":"Expanded Gating Ranges Improve Activation Functions","date":"2024-05-25","arxiv_id":"2405.20768","n_code_links":0,"syntology":null},{"paper":null,"title":"Stable and Robust Deep Learning By Hyperbolic Tangent Exponential Linear Unit (TeLU)","date":"2024-02-05","arxiv_id":"2402.02790","n_code_links":0,"syntology":null},{"paper":"/paper/leveraging-continuously-differentiable","title":"Leveraging Continuously Differentiable Activation Functions for Learning in Quantized Noisy Environments","date":"2024-02-04","arxiv_id":"2402.02593","n_code_links":1,"syntology":null},{"paper":"/paper/relu-strikes-back-exploiting-activation","title":"ReLU Strikes Back: Exploiting Activation Sparsity in Large Language Models","date":"2023-10-06","arxiv_id":"2310.04564","n_code_links":1,"syntology":null},{"paper":"/paper/learnable-extended-activation-function-leaf","title":"Learnable Extended Activation Function (LEAF) for Deep Neural Networks","date":"2023-09-30","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"Attention-Only Transformers and Implementing MLPs with Attention Heads","date":"2023-09-15","arxiv_id":"2309.08593","n_code_links":0,"syntology":null},{"paper":"/paper/compact-approximating-complex-activation","title":"Compact: Approximating Complex Activation Functions for Secure Computation","date":"2023-09-09","arxiv_id":"2309.04664","n_code_links":1,"syntology":null},{"paper":null,"title":"Deep Contract Design via Discontinuous Networks","date":"2023-07-05","arxiv_id":"2307.02318","n_code_links":0,"syntology":null},{"paper":null,"title":"Demystifying Oversmoothing in Attention-Based Graph Neural Networks","date":"2023-05-25","arxiv_id":"2305.16102","n_code_links":0,"syntology":null},{"paper":null,"title":"Saturated Non-Monotonic Activation Functions","date":"2023-05-12","arxiv_id":"2305.07537","n_code_links":0,"syntology":null},{"paper":"/paper/trainable-activations-for-image","title":"Trainable Activations for Image Classification","date":"2023-01-26","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"Optimizing Anchor-based Detectors for Autonomous Driving Scenes","date":"2022-08-11","arxiv_id":"2208.06062","n_code_links":0,"syntology":null},{"paper":"/paper/adaptive-hybrid-activation-function-for-deep","title":"Adaptive hybrid activation function for deep neural networks","date":"2022-04-25","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"A Unified and Constructive Framework for the Universality of Neural Networks","date":"2021-12-30","arxiv_id":"2112.14877","n_code_links":0,"syntology":null},{"paper":"/paper/agglio-global-optimization-for-locally-convex","title":"AGGLIO: Global Optimization for Locally Convex Functions","date":"2021-11-06","arxiv_id":"2111.03932","n_code_links":1,"syntology":null},{"paper":"/paper/simple-training-strategies-and-model-scaling","title":"Simple Training Strategies and Model Scaling for Object Detection","date":"2021-06-30","arxiv_id":"2107.00057","n_code_links":1,"syntology":null},{"paper":"/paper/low-curvature-activations-reduce-overfitting","title":"Low Curvature Activations Reduce Overfitting in Adversarial Training","date":"2021-02-15","arxiv_id":"2102.07861","n_code_links":1,"syntology":{"ran":1,"of":2,"unverified":1,"pointer_only":2}},{"paper":"/paper/bottleneck-transformers-for-visual","title":"Bottleneck Transformers for Visual Recognition","date":"2021-01-27","arxiv_id":"2101.11605","n_code_links":13,"syntology":{"ran":26,"of":49,"unverified":23,"pointer_only":8}},{"paper":"/paper/searching-for-activation-functions","title":"Searching for Activation Functions","date":"2017-10-16","arxiv_id":"1710.05941","n_code_links":22,"syntology":{"ran":6,"of":20,"unverified":14,"pointer_only":3}},{"paper":"/paper/sigmoid-weighted-linear-units-for-neural","title":"Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning","date":"2017-02-10","arxiv_id":"1702.03118","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/image-classification","name":"Image Classification","papers":8},{"task":"/task/image-classification","name":"image-classification","papers":7},{"task":"/task/architecture-search","name":"Neural Architecture Search","papers":3},{"task":"/task/activation-function-synthesis","name":"Activation Function Synthesis","papers":2},{"task":null,"name":"GPU","papers":2},{"task":"/task/instance-segmentation","name":"Instance Segmentation","papers":2},{"task":"/task/learning-theory","name":"Learning Theory","papers":2},{"task":"/task/object-detection","name":"Object Detection","papers":2},{"task":"/task/privacy-preserving","name":"Privacy Preserving","papers":2},{"task":"/task/reinforcement-learning","name":"Reinforcement Learning","papers":2},{"task":"/task/segmentation","name":"Segmentation","papers":2},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":2},{"task":"/task/object-detection-1","name":"object-detection","papers":2},{"task":"/task/arc","name":"ARC","papers":1},{"task":"/task/atari-games","name":"Atari Games","papers":1},{"task":"/task/autonomous-driving","name":"Autonomous Driving","papers":1},{"task":null,"name":"CPU","papers":1},{"task":"/task/computational-efficiency","name":"Computational Efficiency","papers":1},{"task":"/task/deep-learning","name":"Deep Learning","papers":1},{"task":"/task/deep-reinforcement-learning","name":"Deep Reinforcement Learning","papers":1}],"tasks_shown":20,"n_tasks":41,"usage_by_year":[{"year":"2017","papers":2},{"year":"2021","papers":5},{"year":"2022","papers":2},{"year":"2023","papers":8},{"year":"2024","papers":12},{"year":"2025","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/silu"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}