{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/classifying-idiomatic-and-literal-expressions-2","title":"Classifying Idiomatic and Literal Expressions Using Topic Models and Intensity of Emotions","arxiv_id":"1802.09961","date":"2018-02-27","proceeding":"EMNLP 2014 10","authors":["Jing Peng","Anna Feldman","Ekaterina Vylomova"],"abstract":"We describe an algorithm for automatic classification of idiomatic and\nliteral expressions. Our starting point is that words in a given text segment,\nsuch as a paragraph, that are highranking representatives of a common topic of\ndiscussion are less likely to be a part of an idiomatic expression. Our\nadditional hypothesis is that contexts in which idioms occur, typically, are\nmore affective and therefore, we incorporate a simple analysis of the intensity\nof the emotions expressed by the contexts. We investigate the bag of words\ntopic representation of one to three paragraphs containing an expression that\nshould be classified as idiomatic or literal (a target phrase). We extract\ntopics from paragraphs containing idioms and from paragraphs containing\nliterals using an unsupervised clustering method, Latent Dirichlet Allocation\n(LDA) (Blei et al., 2003). Since idiomatic expressions exhibit the property of\nnon-compositionality, we assume that they usually present different semantics\nthan the words used in the local topic. We treat idioms as semantic outliers,\nand the identification of a semantic shift as outlier detection. Thus, this\ntopic representation allows us to differentiate idioms from literals using\nlocal semantic contexts. Our results are encouraging.","url_abs":"http://arxiv.org/abs/1802.09961v1","url_pdf":"http://arxiv.org/pdf/1802.09961v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"classifying-idiomatic-and-literal-expressions-2","repo_url":"https://github.com/bondfeld/BNC_idioms","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"outlier-detection","task_name":"Outlier Detection"},{"task_slug":"topic-models","task_name":"Topic Models"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1802.09961","atlas_url":"https://app.syntology.ai/?focus=1802.09961","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}