{"url":"/method/elmo","slug":"elmo","name":"ELMo","full_name":"ELMo","full_name_withheld":false,"description_markdown":"**Embeddings from Language Models**, or **ELMo**, is a type of deep contextualized word representation that models both (1) complex characteristics of word use (e.g., syntax and semantics), and (2) how these uses vary across linguistic contexts (i.e., to model polysemy). Word vectors are learned functions of the internal states of a deep bidirectional language model (biLM), which is pre-trained on a large text corpus.\r\n\r\nA biLM combines both a forward and backward LM.  ELMo jointly maximizes the log likelihood of the forward and backward directions. To add ELMo to a supervised model, we freeze the weights of the biLM and then concatenate the ELMo vector $\\textbf{ELMO}^{task}_k$ with $\\textbf{x}_k$ and pass the ELMO enhanced representation $[\\textbf{x}_k; \\textbf{ELMO}^{task}_k]$ into the task RNN. Here $\\textbf{x}_k$ is a context-independent token representation for each token position. \r\n\r\nImage Source: [here](https://medium.com/@duyanhnguyen_38925/create-a-strong-text-classification-with-the-help-from-elmo-e90809ba29da)","description_state":"present","introduced_year":null,"introduced_by":{"title":"Deep contextualized word representations","paper":"/paper/deep-contextualized-word-representations","first_author":"Matthew E. Peters","n_authors":7,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/deep-contextualized-word-representations"},"source":{"url":"http://arxiv.org/abs/1802.05365v2","title":"Deep contextualized word representations","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Language Models","url":"/methods/category/language-models","pwc_aliases":[]},{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Contextualized Word Embeddings","url":"/methods/category/contextualized-word-embeddings","pwc_aliases":[]},{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Word Embeddings","url":"/methods/category/word-embeddings","pwc_aliases":[]}],"n_papers_tagged":234,"archive_num_papers":234,"papers_newest_first":[{"paper":"/paper/a-comparative-analysis-of-static-word","title":"A Comparative Analysis of Static Word Embeddings for Hungarian","date":"2025-05-12","arxiv_id":"2505.07809","n_code_links":1,"syntology":null},{"paper":null,"title":"Embedding-based Approaches to Hyperpartisan News Detection","date":"2025-01-02","arxiv_id":"2501.01370","n_code_links":0,"syntology":null},{"paper":"/paper/generative-pretrained-embedding-and","title":"Generative Pretrained Embedding and Hierarchical Irregular Time Series Representation for Daily Living Activity Recognition","date":"2024-12-27","arxiv_id":"2412.19732","n_code_links":1,"syntology":null},{"paper":null,"title":"LLMs are Also Effective Embedding Models: An In-depth Overview","date":"2024-12-17","arxiv_id":"2412.12591","n_code_links":0,"syntology":null},{"paper":null,"title":"From Word Vectors to Multimodal Embeddings: Techniques, Applications, and Future Directions For Large Language Models","date":"2024-11-06","arxiv_id":"2411.05036","n_code_links":0,"syntology":null},{"paper":null,"title":"ELMO: Enhanced Real-time LiDAR Motion Capture through Upsampling","date":"2024-10-09","arxiv_id":"2410.06963","n_code_links":0,"syntology":null},{"paper":null,"title":"Evaluating the Efficacy of AI Techniques in Textual Anonymization: A Comparative Study","date":"2024-05-09","arxiv_id":"2405.06709","n_code_links":0,"syntology":null},{"paper":null,"title":"Where exactly does contextualization in a PLM happen?","date":"2023-12-11","arxiv_id":"2312.06514","n_code_links":0,"syntology":null},{"paper":"/paper/semantic-change-detection-for-the-romanian","title":"Semantic Change Detection for the Romanian Language","date":"2023-08-23","arxiv_id":"2308.12131","n_code_links":1,"syntology":null},{"paper":"/paper/pevolm-protein-sequence-evolutionary","title":"PEvoLM: Protein Sequence Evolutionary Information Language Model","date":"2023-08-16","arxiv_id":"2308.08578","n_code_links":1,"syntology":null},{"paper":null,"title":"On \"Scientific Debt\" in NLP: A Case for More Rigour in Language Model Pre-Training Research","date":"2023-06-05","arxiv_id":"2306.02870","n_code_links":0,"syntology":null},{"paper":null,"title":"Analyzing the Generalizability of Deep Contextualized Language Representations For Text Classification","date":"2023-03-22","arxiv_id":"2303.12936","n_code_links":0,"syntology":null},{"paper":null,"title":"Classifying Text-Based Conspiracy Tweets related to COVID-19 using Contextualized Word Embeddings","date":"2023-03-07","arxiv_id":"2303.03706","n_code_links":0,"syntology":null},{"paper":null,"title":"CKG: Dynamic Representation Based on Context and Knowledge Graph","date":"2022-12-09","arxiv_id":"2212.04909","n_code_links":0,"syntology":null},{"paper":null,"title":"A Context-Sensitive Word Embedding Approach for The Detection of Troll Tweets","date":"2022-07-17","arxiv_id":"2207.08230","n_code_links":0,"syntology":null},{"paper":"/paper/always-keep-your-target-in-mind-studying-1","title":"Always Keep your Target in Mind: Studying Semantics and Improving Performance of Neural Lexical Substitution","date":"2022-06-07","arxiv_id":"2206.11815","n_code_links":1,"syntology":null},{"paper":null,"title":"Parameter-Efficient Tuning by Manipulating Hidden States of Pretrained Language Models For Classification Tasks","date":"2022-04-10","arxiv_id":"2204.04596","n_code_links":0,"syntology":null},{"paper":"/paper/bridging-pre-trained-language-models-and-hand","title":"Bridging Pre-trained Language Models and Hand-crafted Features for Unsupervised POS Tagging","date":"2022-03-19","arxiv_id":"2203.10315","n_code_links":1,"syntology":null},{"paper":null,"title":"Using Word Embeddings to Analyze Protests News","date":"2022-03-11","arxiv_id":"2203.05875","n_code_links":0,"syntology":null},{"paper":null,"title":"Assessment of contextualised representations in detecting outcome phrases in clinical trials","date":"2022-02-13","arxiv_id":"2203.03547","n_code_links":0,"syntology":null},{"paper":"/paper/a-passage-to-india-pre-trained-word-1","title":"\"A Passage to India\": Pre-trained Word Embeddings for Indian Languages","date":"2021-12-27","arxiv_id":"2112.13800","n_code_links":0,"syntology":null},{"paper":null,"title":"KARL-Trans-NER: Knowledge Aware Representation Learning for Named Entity Recognition using Transformers","date":"2021-11-30","arxiv_id":"2111.15436","n_code_links":0,"syntology":null},{"paper":"/paper/using-language-model-to-bootstrap-human","title":"Using Language Model to Bootstrap Human Activity Recognition Ambient Sensors Based in Smart Homes","date":"2021-11-23","arxiv_id":"2111.12158","n_code_links":1,"syntology":null},{"paper":null,"title":"Investigating the Use of BERT Anchors for Bilingual Lexicon Induction with Minimal Supervision","date":"2021-11-16","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":"/paper/fake-news-detection-in-spanish-using-deep","title":"Fake News Detection in Spanish Using Deep Learning Techniques","date":"2021-10-13","arxiv_id":"2110.06461","n_code_links":1,"syntology":null},{"paper":"/paper/a-comprehensive-comparison-of-word-embeddings","title":"A Comprehensive Comparison of Word Embeddings in Event & Entity Coreference Resolution","date":"2021-10-11","arxiv_id":"2110.05115","n_code_links":1,"syntology":null},{"paper":null,"title":"Cross-Architecture Distillation Using Bidirectional CMOW Embeddings","date":"2021-09-29","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":"/paper/new-students-on-sesame-street-what-order","title":"General Cross-Architecture Distillation of Pretrained Language Models into Matrix Embeddings","date":"2021-09-17","arxiv_id":"2109.08449","n_code_links":1,"syntology":null},{"paper":"/paper/revisiting-tri-training-of-dependency-parsers","title":"Revisiting Tri-training of Dependency Parsers","date":"2021-09-16","arxiv_id":"2109.08122","n_code_links":2,"syntology":null},{"paper":null,"title":"Sense representations for Portuguese: experiments with sense embeddings and deep neural language models","date":"2021-08-31","arxiv_id":"2109.00025","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/word-embeddings","name":"Word Embeddings","papers":71},{"task":"/task/sentence","name":"Sentence","papers":48},{"task":"/task/language-modelling","name":"Language Modelling","papers":38},{"task":"/task/language-modeling","name":"Language Modeling","papers":33},{"task":"/task/sentiment-analysis","name":"Sentiment Analysis","papers":26},{"task":"/task/named-entity-recognition-ner","name":"Named Entity Recognition (NER)","papers":22},{"task":"/task/cg","name":"NER","papers":20},{"task":"/task/named-entity-recognition-1","name":"Named Entity Recognition","papers":20},{"task":"/task/named-entity-recognition","name":"named-entity-recognition","papers":20},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":18},{"task":"/task/text-classification-1","name":"text-classification","papers":16},{"task":"/task/classification","name":"General Classification","papers":15},{"task":"/task/text-classification","name":"Text Classification","papers":15},{"task":"/task/question-answering","name":"Question Answering","papers":14},{"task":"/task/natural-language-inference","name":"Natural Language Inference","papers":12},{"task":"/task/pos","name":"POS","papers":11},{"task":"/task/word-sense-disambiguation","name":"Word Sense Disambiguation","papers":11},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":10},{"task":"/task/dependency-parsing","name":"Dependency Parsing","papers":8},{"task":"/task/machine-translation","name":"Machine Translation","papers":8}],"tasks_shown":20,"n_tasks":199,"usage_by_year":[{"year":"2018","papers":25},{"year":"2019","papers":94},{"year":"2020","papers":59},{"year":"2021","papers":36},{"year":"2022","papers":7},{"year":"2023","papers":6},{"year":"2024","papers":5},{"year":"2025","papers":2}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/elmo"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}