{"url":"/method/lstm","slug":"lstm","name":"LSTM","full_name":"Long Short-Term Memory","full_name_withheld":false,"description_markdown":"An **LSTM** is a type of [recurrent neural network](https://paperswithcode.com/methods/category/recurrent-neural-networks) that addresses the vanishing gradient problem in vanilla RNNs through additional cells, input and output gates. Intuitively, vanishing gradients are solved through additional *additive* components, and forget gate activations, that allow the gradients to flow through the network without vanishing as quickly.\r\n\r\n(Image Source [here](https://medium.com/datadriveninvestor/how-do-lstm-networks-solve-the-problem-of-vanishing-gradients-a6784971a577))\r\n\r\n(Introduced by Hochreiter and Schmidhuber)","description_state":"present","introduced_year":1997,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":null,"title":null,"url_on_a_paper_host":false},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Sequential","area_id":"sequential","collection":"Recurrent Neural Networks","url":"/methods/category/recurrent-neural-networks","pwc_aliases":[]}],"n_papers_tagged":5448,"archive_num_papers":5448,"papers_newest_first":[{"paper":null,"title":"AI-Based Demand Forecasting and Load Balancing for Optimising Energy use in Healthcare Systems: A real case study","date":"2025-07-08","arxiv_id":"2507.06077","n_code_links":0,"syntology":null},{"paper":null,"title":"A Hybrid Machine Learning Framework for Optimizing Crop Selection via Agronomic and Economic Forecasting","date":"2025-07-06","arxiv_id":"2507.08832","n_code_links":0,"syntology":null},{"paper":"/paper/mambattention-mamba-with-multi-head-attention","title":"MambAttention: Mamba with Multi-Head Attention for Generalizable Single-Channel Speech Enhancement","date":"2025-07-01","arxiv_id":"2507.00966","n_code_links":2,"syntology":null},{"paper":null,"title":"Dense Video Captioning using Graph-based Sentence Summarization","date":"2025-06-25","arxiv_id":"2506.20583","n_code_links":0,"syntology":null},{"paper":null,"title":"Efficacy of Temporal Fusion Transformers for Runoff Simulation","date":"2025-06-25","arxiv_id":"2506.20831","n_code_links":0,"syntology":null},{"paper":null,"title":"FINN-GL: Generalized Mixed-Precision Extensions for FPGA-Accelerated LSTMs","date":"2025-06-25","arxiv_id":"2506.20810","n_code_links":0,"syntology":null},{"paper":null,"title":"Show, Tell and Summarize: Dense Video Captioning Using Visual Cue Aided Sentence Summarization","date":"2025-06-25","arxiv_id":"2506.20567","n_code_links":0,"syntology":null},{"paper":"/paper/cove-compressed-vocabulary-expansion-makes","title":"CoVE: Compressed Vocabulary Expansion Makes Better LLM-based Recommender Systems","date":"2025-06-24","arxiv_id":"2506.19993","n_code_links":1,"syntology":{"ran":0,"of":4,"unverified":4,"pointer_only":4}},{"paper":null,"title":"Emotion Detection on User Front-Facing App Interfaces for Enhanced Schedule Optimization: A Machine Learning Approach","date":"2025-06-24","arxiv_id":"2506.19280","n_code_links":0,"syntology":null},{"paper":null,"title":"Simulation of a closed-loop dc-dc converter using a physics-informed neural network-based model","date":"2025-06-23","arxiv_id":"2506.19178","n_code_links":0,"syntology":null},{"paper":null,"title":"Programmable-Room: Interactive Textured 3D Room Meshes Generation Empowered by Large Language Models","date":"2025-06-21","arxiv_id":"2506.17707","n_code_links":0,"syntology":null},{"paper":null,"title":"Exploring Speaker Diarization with Mixture of Experts","date":"2025-06-17","arxiv_id":"2506.14750","n_code_links":0,"syntology":null},{"paper":null,"title":"Intelligent Image Sensing for Crime Analysis: A ML Approach towards Enhanced Violence Detection and Investigation","date":"2025-06-16","arxiv_id":"2506.13910","n_code_links":0,"syntology":null},{"paper":null,"title":"Seq2Bind Webserver for Decoding Binding Hotspots directly from Sequences using Fine-Tuned Protein Language Models","date":"2025-06-16","arxiv_id":"2506.13830","n_code_links":0,"syntology":null},{"paper":null,"title":"Transforming Chatbot Text: A Sequence-to-Sequence Approach","date":"2025-06-15","arxiv_id":"2506.12843","n_code_links":0,"syntology":null},{"paper":null,"title":"Brain2Vec: A Deep Learning Framework for EEG-Based Stress Detection Using CNN-LSTM-Attention","date":"2025-06-12","arxiv_id":"2506.11179","n_code_links":0,"syntology":null},{"paper":null,"title":"Data-driven Day Ahead Market Prices Forecasting: A Focus on Short Training Set Windows","date":"2025-06-12","arxiv_id":"2506.10536","n_code_links":0,"syntology":null},{"paper":null,"title":"Improving the performance of optical inverse design of multilayer thin films using CNN-LSTM tandem neural networks","date":"2025-06-11","arxiv_id":"2506.10044","n_code_links":0,"syntology":null},{"paper":null,"title":"Cross-Learning Between ECG and PCG: Exploring Common and Exclusive Characteristics of Bimodal Electromechanical Cardiac Waveforms","date":"2025-06-11","arxiv_id":"2506.10212","n_code_links":0,"syntology":null},{"paper":null,"title":"Analyzing Emotions in Bangla Social Media Comments Using Machine Learning and LIME","date":"2025-06-11","arxiv_id":"2506.10154","n_code_links":0,"syntology":null},{"paper":null,"title":"Enhancing Bagging Ensemble Regression with Data Integration for Time Series-Based Diabetes Prediction","date":"2025-06-11","arxiv_id":"2506.13786","n_code_links":0,"syntology":null},{"paper":null,"title":"Data Augmentation For Small Object using Fast AutoAugment","date":"2025-06-10","arxiv_id":"2506.08956","n_code_links":0,"syntology":null},{"paper":null,"title":"Hyperpruning: Efficient Search through Pruned Variants of Recurrent Neural Networks Leveraging Lyapunov Spectrum","date":"2025-06-09","arxiv_id":"2506.07975","n_code_links":0,"syntology":null},{"paper":null,"title":"A Statistical Framework for Model Selection in LSTM Networks","date":"2025-06-07","arxiv_id":"2506.06840","n_code_links":0,"syntology":null},{"paper":"/paper/information-locality-as-an-inductive-bias-for","title":"Information Locality as an Inductive Bias for Neural Language Models","date":"2025-06-05","arxiv_id":"2506.05136","n_code_links":1,"syntology":null},{"paper":null,"title":"Sentiment Analysis in Learning Management Systems Understanding Student Feedback at Scale","date":"2025-06-05","arxiv_id":"2506.05490","n_code_links":0,"syntology":null},{"paper":null,"title":"Frame-Level Real-Time Assessment of Stroke Rehabilitation Exercises from Video-Level Labeled Data: Task-Specific vs. Foundation Models","date":"2025-06-04","arxiv_id":"2506.03752","n_code_links":0,"syntology":null},{"paper":null,"title":"Heart Rate Classification in ECG Signals Using Machine Learning and Deep Learning","date":"2025-06-02","arxiv_id":"2506.06349","n_code_links":0,"syntology":null},{"paper":"/paper/cnn-lstm-hybrid-model-for-ai-driven","title":"CNN-LSTM Hybrid Model for AI-Driven Prediction of COVID-19 Severity from Spike Sequences and Clinical Data","date":"2025-05-29","arxiv_id":"2505.23879","n_code_links":1,"syntology":null},{"paper":null,"title":"Combining Deep Architectures for Information Gain estimation and Reinforcement Learning for multiagent field exploration","date":"2025-05-29","arxiv_id":"2505.23865","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/sentence","name":"Sentence","papers":427},{"task":"/task/time-series-1","name":"Time Series","papers":398},{"task":"/task/language-modelling","name":"Language Modelling","papers":390},{"task":"/task/time-series","name":"Time Series Analysis","papers":334},{"task":"/task/decoder","name":"Decoder","papers":330},{"task":"/task/language-modeling","name":"Language Modeling","papers":318},{"task":"/task/classification","name":"General Classification","papers":301},{"task":"/task/architecture-search","name":"Neural Architecture Search","papers":273},{"task":"/task/prediction","name":"Prediction","papers":258},{"task":"/task/sentiment-analysis","name":"Sentiment Analysis","papers":258},{"task":"/task/machine-translation","name":"Machine Translation","papers":226},{"task":"/task/word-embeddings","name":"Word Embeddings","papers":209},{"task":"/task/translation","name":"Translation","papers":204},{"task":"/task/classification-1","name":"Classification","papers":200},{"task":"/task/deep-learning","name":"Deep Learning","papers":200},{"task":"/task/speech-recognition","name":"Speech Recognition","papers":193},{"task":"/task/speech-recognition-1","name":"speech-recognition","papers":183},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":159},{"task":"/task/reinforcement-learning","name":"Reinforcement Learning","papers":151},{"task":"/task/text-classification","name":"Text Classification","papers":139}],"tasks_shown":20,"n_tasks":1283,"usage_by_year":[{"year":"2014","papers":13},{"year":"2015","papers":56},{"year":"2016","papers":175},{"year":"2017","papers":316},{"year":"2018","papers":612},{"year":"2019","papers":923},{"year":"2020","papers":876},{"year":"2021","papers":750},{"year":"2022","papers":499},{"year":"2023","papers":502},{"year":"2024","papers":491},{"year":"2025","papers":235}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/lstm"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}