{"url":"/method/tanh-activation","slug":"tanh-activation","name":"Tanh Activation","full_name":"Tanh Activation","full_name_withheld":false,"description_markdown":"**Tanh Activation** is an activation function used for neural networks:\r\n\r\n$$f\\left(x\\right) = \\frac{e^{x} - e^{-x}}{e^{x} + e^{-x}}$$\r\n\r\nHistorically, the tanh function became preferred over the [sigmoid function](https://paperswithcode.com/method/sigmoid-activation) as it gave better performance for multi-layer neural networks. But it did not solve the vanishing gradient problem that sigmoids suffered, which was tackled more effectively with the introduction of [ReLU](https://paperswithcode.com/method/relu) activations.\r\n\r\nImage Source: [Junxi Feng](https://www.researchgate.net/profile/Junxi_Feng)","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":null,"title":null,"url_on_a_paper_host":false},"code_snippet_url":"https://github.com/pytorch/pytorch/blob/96aaa311c0251d24decb9dc5da4957b7c590af6f/torch/nn/modules/activation.py#L329","code_snippet_url_on_a_code_host":true,"categories":[{"area":"General","area_id":"general","collection":"Activation Functions","url":"/methods/category/activation-functions","pwc_aliases":[]}],"n_papers_tagged":6333,"archive_num_papers":null,"papers_newest_first":[{"paper":"/paper/dual-detector-re-optimization-for-federated","title":"Dual‑detector Re‑optimization for Federated Weakly Supervised Video Anomaly Detection Via Adaptive Dynamic Recursive Mapping","date":"2025-06-13","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":null,"title":"Hyperpruning: Efficient Search through Pruned Variants of Recurrent Neural Networks Leveraging Lyapunov Spectrum","date":"2025-06-09","arxiv_id":"2506.07975","n_code_links":0,"syntology":null},{"paper":null,"title":"Deep Learning Weather Models for Subregional Ocean Forecasting: A Case Study on the Canary Current Upwelling System","date":"2025-05-30","arxiv_id":"2505.24429","n_code_links":0,"syntology":null},{"paper":"/paper/cnn-lstm-hybrid-model-for-ai-driven","title":"CNN-LSTM Hybrid Model for AI-Driven Prediction of COVID-19 Severity from Spike Sequences and Clinical Data","date":"2025-05-29","arxiv_id":"2505.23879","n_code_links":1,"syntology":null},{"paper":null,"title":"Combining Deep Architectures for Information Gain estimation and Reinforcement Learning for multiagent field exploration","date":"2025-05-29","arxiv_id":"2505.23865","n_code_links":0,"syntology":null},{"paper":null,"title":"Gradient Boosting Decision Tree with LSTM for Investment Prediction","date":"2025-05-29","arxiv_id":"2505.23084","n_code_links":0,"syntology":null},{"paper":"/paper/tirex-zero-shot-forecasting-across-long-and","title":"TiRex: Zero-Shot Forecasting Across Long and Short Horizons with Enhanced In-Context Learning","date":"2025-05-29","arxiv_id":"2505.23719","n_code_links":1,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":1}},{"paper":null,"title":"Multipath cycleGAN for harmonization of paired and unpaired low-dose lung computed tomography reconstruction kernels","date":"2025-05-28","arxiv_id":"2505.22568","n_code_links":0,"syntology":null},{"paper":"/paper/a-domain-adaptation-neural-network-for","title":"A domain adaptation neural network for digital twin-supported fault diagnosis","date":"2025-05-27","arxiv_id":"2505.21046","n_code_links":1,"syntology":null},{"paper":"/paper/localized-weather-prediction-using-kolmogorov","title":"Localized Weather Prediction Using Kolmogorov-Arnold Network-Based Models and Deep RNNs","date":"2025-05-27","arxiv_id":"2505.22686","n_code_links":1,"syntology":null},{"paper":null,"title":"Unpaired Image-to-Image Translation for Segmentation and Signal Unmixing","date":"2025-05-27","arxiv_id":"2505.20746","n_code_links":0,"syntology":null},{"paper":null,"title":"Gradient Flow Matching for Learning Update Dynamics in Neural Network Training","date":"2025-05-26","arxiv_id":"2505.20221","n_code_links":0,"syntology":null},{"paper":null,"title":"Hybrid Models for Financial Forecasting: Combining Econometric, Machine Learning, and Deep Learning Models","date":"2025-05-26","arxiv_id":"2505.19617","n_code_links":0,"syntology":null},{"paper":null,"title":"Multiple Descents in Deep Learning as a Sequence of Order-Chaos Transitions","date":"2025-05-26","arxiv_id":"2505.20030","n_code_links":0,"syntology":null},{"paper":null,"title":"PosePilot: An Edge-AI Solution for Posture Correction in Physical Exercises","date":"2025-05-25","arxiv_id":"2505.19186","n_code_links":0,"syntology":null},{"paper":null,"title":"SETransformer: A Hybrid Attention-Based Architecture for Robust Human Activity Recognition","date":"2025-05-25","arxiv_id":"2505.19369","n_code_links":0,"syntology":null},{"paper":null,"title":"Improving Bangla Linguistics: Advanced LSTM, Bi-LSTM, and Seq2Seq Models for Translating Sylheti to Modern Bangla","date":"2025-05-24","arxiv_id":"2505.18709","n_code_links":0,"syntology":null},{"paper":null,"title":"Is Attention Required for Transformer Inference? Explore Function-preserving Attention Replacement","date":"2025-05-24","arxiv_id":"2505.21535","n_code_links":0,"syntology":null},{"paper":null,"title":"Pruning for Performance: Efficient Idiom and Metaphor Classification in Low-Resource Konkani Using mBERT","date":"2025-05-24","arxiv_id":"2506.02005","n_code_links":0,"syntology":null},{"paper":null,"title":"Riverine Flood Prediction and Early Warning in Mountainous Regions using Artificial Intelligence","date":"2025-05-24","arxiv_id":"2505.18645","n_code_links":0,"syntology":null},{"paper":null,"title":"Enhancing AI System Resiliency: Formulation and Guarantee for LSTM Resilience Based on Control Theory","date":"2025-05-23","arxiv_id":"2505.17696","n_code_links":0,"syntology":null},{"paper":null,"title":"Towards a Quantum-classical Augmented Network","date":"2025-05-23","arxiv_id":"2505.18282","n_code_links":0,"syntology":null},{"paper":null,"title":"Tube Loss based Deep Networks For Improving the Probabilistic Forecasting of Wind Speed","date":"2025-05-23","arxiv_id":"2505.18284","n_code_links":0,"syntology":null},{"paper":null,"title":"End-to-End Framework for Predicting the Remaining Useful Life of Lithium-Ion Batteries","date":"2025-05-22","arxiv_id":"2505.16664","n_code_links":0,"syntology":null},{"paper":null,"title":"NY Real Estate Racial Equity Analysis via Applied Machine Learning","date":"2025-05-22","arxiv_id":"2505.16946","n_code_links":0,"syntology":null},{"paper":null,"title":"Convolutional Long Short-Term Memory Neural Networks Based Numerical Simulation of Flow Field","date":"2025-05-21","arxiv_id":"2505.15533","n_code_links":0,"syntology":null},{"paper":null,"title":"EEG-Based Inter-Patient Epileptic Seizure Detection Combining Domain Adversarial Training with CNN-BiLSTM Network","date":"2025-05-21","arxiv_id":"2505.15203","n_code_links":0,"syntology":null},{"paper":null,"title":"Machine Learning Derived Blood Input for Dynamic PET Images of Rat Heart","date":"2025-05-21","arxiv_id":"2505.15488","n_code_links":0,"syntology":null},{"paper":"/paper/rlbenchnet-the-right-network-for-the-right","title":"RLBenchNet: The Right Network for the Right Reinforcement Learning Task","date":"2025-05-21","arxiv_id":"2505.15040","n_code_links":1,"syntology":null},{"paper":null,"title":"3D Reconstruction from Sketches","date":"2025-05-20","arxiv_id":"2505.14621","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/sentence","name":"Sentence","papers":436},{"task":"/task/time-series-1","name":"Time Series","papers":406},{"task":"/task/language-modelling","name":"Language Modelling","papers":405},{"task":"/task/decoder","name":"Decoder","papers":358},{"task":"/task/time-series","name":"Time Series Analysis","papers":340},{"task":"/task/language-modeling","name":"Language Modeling","papers":332},{"task":"/task/classification","name":"General Classification","papers":313},{"task":"/task/translation","name":"Translation","papers":312},{"task":"/task/architecture-search","name":"Neural Architecture Search","papers":277},{"task":"/task/prediction","name":"Prediction","papers":276},{"task":"/task/sentiment-analysis","name":"Sentiment Analysis","papers":258},{"task":"/task/deep-learning","name":"Deep Learning","papers":235},{"task":"/task/machine-translation","name":"Machine Translation","papers":231},{"task":"/task/classification-1","name":"Classification","papers":216},{"task":"/task/word-embeddings","name":"Word Embeddings","papers":209},{"task":"/task/speech-recognition","name":"Speech Recognition","papers":208},{"task":"/task/speech-recognition-1","name":"speech-recognition","papers":196},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":186},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":162},{"task":"/task/reinforcement-learning","name":"Reinforcement Learning","papers":155}],"tasks_shown":20,"n_tasks":1433,"usage_by_year":[{"year":"2014","papers":14},{"year":"2015","papers":57},{"year":"2016","papers":179},{"year":"2017","papers":332},{"year":"2018","papers":659},{"year":"2019","papers":1009},{"year":"2020","papers":1098},{"year":"2021","papers":898},{"year":"2022","papers":643},{"year":"2023","papers":609},{"year":"2024","papers":583},{"year":"2025","papers":252}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/tanh-activation"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}