{"url":"/method/elu","slug":"elu","name":"ELU","full_name":"Exponential Linear Unit","full_name_withheld":false,"description_markdown":"The **Exponential Linear Unit** (ELU) is an activation function for neural networks. In contrast to [ReLUs](https://paperswithcode.com/method/relu), ELUs have negative values which allows them to push mean unit activations closer to zero like [batch normalization](https://paperswithcode.com/method/batch-normalization) but with lower computational complexity. Mean shifts toward zero speed up learning by bringing the normal gradient closer to the unit natural gradient because of a reduced bias shift effect. While [LReLUs](https://paperswithcode.com/method/leaky-relu) and [PReLUs](https://paperswithcode.com/method/prelu) have negative values, too, they do not ensure a noise-robust deactivation state. ELUs saturate to a negative value with smaller inputs and thereby decrease the forward propagated variation and information.\r\n\r\nThe exponential linear unit (ELU) with $0 < \\alpha$ is:\r\n\r\n$$f\\left(x\\right) = x \\text{ if } x > 0$$\r\n$$\\alpha\\left(\\exp\\left(x\\right) − 1\\right) \\text{ if } x \\leq 0$$","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"http://arxiv.org/abs/1511.07289v5","title":"Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/pytorch/pytorch/blob/96aaa311c0251d24decb9dc5da4957b7c590af6f/torch/nn/modules/activation.py#L422","code_snippet_url_on_a_code_host":true,"categories":[{"area":"General","area_id":"general","collection":"Activation Functions","url":"/methods/category/activation-functions","pwc_aliases":[]}],"n_papers_tagged":44,"archive_num_papers":null,"papers_newest_first":[{"paper":"/paper/pause-low-latency-and-privacy-aware-active","title":"PAUSE: Low-Latency and Privacy-Aware Active User Selection for Federated Learning","date":"2025-03-17","arxiv_id":"2503.13173","n_code_links":1,"syntology":null},{"paper":"/paper/evaluating-the-performance-of-taaf-for-image","title":"Evaluating the Performance of TAAF for image classification models","date":"2025-02-13","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/gompertz-linear-units-leveraging-asymmetry-1","title":"Gompertz Linear Units: Leveraging Asymmetry for Enhanced Learning Dynamics","date":"2025-02-05","arxiv_id":"2502.03654","n_code_links":1,"syntology":null},{"paper":"/paper/deriving-activation-functions-via-integration","title":"Deriving Activation Functions Using Integration","date":"2024-11-20","arxiv_id":"2411.13010","n_code_links":1,"syntology":null},{"paper":"/paper/elu-gcn-effectively-label-utilizing-graph","title":"Enhancing the Influence of Labels on Unlabeled Nodes in Graph Convolutional Networks","date":"2024-11-04","arxiv_id":"2411.02279","n_code_links":1,"syntology":null},{"paper":null,"title":"On Expressive Power of Quantized Neural Networks under Fixed-Point Arithmetic","date":"2024-08-30","arxiv_id":"2409.00297","n_code_links":0,"syntology":null},{"paper":null,"title":"SwishReLU: A Unified Approach to Activation Functions for Enhanced Deep Neural Networks Performance","date":"2024-07-11","arxiv_id":"2407.08232","n_code_links":0,"syntology":null},{"paper":null,"title":"Deep Multi-Task Learning for Malware Image Classification","date":"2024-05-09","arxiv_id":"2405.05906","n_code_links":0,"syntology":null},{"paper":"/paper/talu-a-hybrid-activation-function-combining","title":"TaLU: A Hybrid Activation Function Combining Tanh and Rectified Linear Unit to Enhance Neural Networks","date":"2023-05-08","arxiv_id":"2305.04402","n_code_links":1,"syntology":null},{"paper":null,"title":"Ensemble Learning Model on Artificial Neural Network-Backpropagation (ANN-BP) Architecture for Coal Pillar Stability Classification","date":"2023-03-29","arxiv_id":"2303.16524","n_code_links":0,"syntology":null},{"paper":"/paper/lmec-learnable-multiplicative-absolute","title":"LMEC: Learnable Multiplicative Absolute Position Embedding Based Conformer for Speech Recognition","date":"2022-12-05","arxiv_id":"2212.02099","n_code_links":1,"syntology":null},{"paper":"/paper/a-new-activation-for-neural-networks-and-its","title":"SignReLU neural network and its approximation ability","date":"2022-10-19","arxiv_id":"2210.10264","n_code_links":1,"syntology":null},{"paper":null,"title":"TEDL: A Two-stage Evidential Deep Learning Method for Classification Uncertainty Quantification","date":"2022-09-12","arxiv_id":"2209.05522","n_code_links":0,"syntology":null},{"paper":"/paper/constrained-monotonic-neural-networks","title":"Constrained Monotonic Neural Networks","date":"2022-05-24","arxiv_id":"2205.11775","n_code_links":2,"syntology":null},{"paper":null,"title":"Beyond EM Algorithm on Over-specified Two-Component Location-Scale Gaussian Mixtures","date":"2022-05-23","arxiv_id":"2205.11078","n_code_links":0,"syntology":null},{"paper":null,"title":"A Unified and Constructive Framework for the Universality of Neural Networks","date":"2021-12-30","arxiv_id":"2112.14877","n_code_links":0,"syntology":null},{"paper":"/paper/reachability-analysis-of-neural-networks","title":"Reachability analysis of neural networks using mixed monotonicity","date":"2021-11-15","arxiv_id":"2111.07683","n_code_links":1,"syntology":null},{"paper":"/paper/a-comprehensive-survey-and-performance","title":"Activation Functions in Deep Learning: A Comprehensive Survey and Benchmark","date":"2021-09-29","arxiv_id":"2109.14545","n_code_links":1,"syntology":null},{"paper":"/paper/pause-positive-and-annealed-unlabeled","title":"PAUSE: Positive and Annealed Unlabeled Sentence Embedding","date":"2021-09-07","arxiv_id":"2109.03155","n_code_links":1,"syntology":{"ran":2,"of":2,"unverified":0,"pointer_only":0}},{"paper":null,"title":"Research on Optimization Method of Multi-scale Fish Target Fast Detection Network","date":"2021-04-11","arxiv_id":"2104.05050","n_code_links":0,"syntology":null},{"paper":"/paper/identifying-co-adaptation-of-algorithmic-and","title":"Co-Adaptation of Algorithmic and Implementational Innovations in Inference-based Deep Reinforcement Learning","date":"2021-03-31","arxiv_id":"2103.17258","n_code_links":1,"syntology":null},{"paper":null,"title":"Generalized Universal Approximation for Certified Networks","date":"2021-01-01","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":null,"title":"TeLU: A New Activation Function for Deep Learning","date":"2021-01-01","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":null,"title":"Comparisons among different stochastic selection of activation layers for convolutional neural networks for healthcare","date":"2020-11-24","arxiv_id":"2011.11834","n_code_links":0,"syntology":null},{"paper":"/paper/deep-multi-task-learning-for-ssvep-detection","title":"Deep Multi-Task Learning for SSVEP Detection and Visual Response Mapping","date":"2020-10-10","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/groove2groove-one-shot-music-style-transfer","title":"Groove2Groove: One-Shot Music Style Transfer with Supervision from Synthetic Data","date":"2020-08-26","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/reconstruction-bottlenecks-in-object-centric","title":"Reconstruction Bottlenecks in Object-Centric Generative Models","date":"2020-07-13","arxiv_id":"2007.06245","n_code_links":1,"syntology":null},{"paper":null,"title":"Interval Universal Approximation for Neural Networks","date":"2020-07-12","arxiv_id":"2007.06093","n_code_links":0,"syntology":null},{"paper":null,"title":"Overcoming Overfitting and Large Weight Update Problem in Linear Rectifiers: Thresholded Exponential Rectified Linear Units","date":"2020-06-04","arxiv_id":"2006.02797","n_code_links":0,"syntology":null},{"paper":"/paper/spleeter-a-fast-and-state-of-the-art-music","title":"Spleeter: A Fast And State-of-the Art Music Source Separation Tool With Pre-trained Models","date":"2019-11-04","arxiv_id":null,"n_code_links":3,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/deep-learning","name":"Deep Learning","papers":5},{"task":"/task/image-classification","name":"Image Classification","papers":4},{"task":"/task/classification","name":"General Classification","papers":3},{"task":"/task/image-classification","name":"image-classification","papers":3},{"task":"/task/classification-1","name":"Classification","papers":2},{"task":"/task/decoder","name":"Decoder","papers":2},{"task":"/task/multi-task-learning","name":"Multi-Task Learning","papers":2},{"task":"/task/object","name":"Object","papers":2},{"task":"/task/object-discovery","name":"Object Discovery","papers":2},{"task":"/task/representation-learning","name":"Representation Learning","papers":2},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":2},{"task":"/task/computational-efficiency","name":"Computational Efficiency","papers":1},{"task":"/task/contrastive-learning","name":"Contrastive Learning","papers":1},{"task":"/task/deep-reinforcement-learning","name":"Deep Reinforcement Learning","papers":1},{"task":"/task/denoising","name":"Denoising","papers":1},{"task":"/task/eeg","name":"Electroencephalogram (EEG)","papers":1},{"task":"/task/ensemble-learning","name":"Ensemble Learning","papers":1},{"task":"/task/federated-learning","name":"Federated Learning","papers":1},{"task":null,"name":"GPU","papers":1},{"task":"/task/graph-learning","name":"Graph Learning","papers":1}],"tasks_shown":20,"n_tasks":59,"usage_by_year":[{"year":"2015","papers":1},{"year":"2016","papers":2},{"year":"2017","papers":4},{"year":"2018","papers":4},{"year":"2019","papers":4},{"year":"2020","papers":6},{"year":"2021","papers":8},{"year":"2022","papers":5},{"year":"2023","papers":2},{"year":"2024","papers":5},{"year":"2025","papers":3}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/elu"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}