{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/calculating-the-similarity-between-words-and","title":"Calculating the similarity between words and sentences using a lexical database and corpus statistics","arxiv_id":"1802.05667","date":"2018-02-15","proceeding":null,"authors":["Atish Pawar","Vijay Mago"],"abstract":"Calculating the semantic similarity between sentences is a long dealt problem\nin the area of natural language processing. The semantic analysis field has a\ncrucial role to play in the research related to the text analytics. The\nsemantic similarity differs as the domain of operation differs. In this paper,\nwe present a methodology which deals with this issue by incorporating semantic\nsimilarity and corpus statistics. To calculate the semantic similarity between\nwords and sentences, the proposed method follows an edge-based approach using a\nlexical database. The methodology can be applied in a variety of domains. The\nmethodology has been tested on both benchmark standards and mean human\nsimilarity dataset. When tested on these two datasets, it gives highest\ncorrelation value for both word and sentence similarity outperforming other\nsimilar models. For word similarity, we obtained Pearson correlation\ncoefficient of 0.8753 and for sentence similarity, the correlation obtained is\n0.8794.","url_abs":"http://arxiv.org/abs/1802.05667v2","url_pdf":"http://arxiv.org/pdf/1802.05667v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"calculating-the-similarity-between-words-and","repo_url":"https://github.com/IljaSamoilov/subtitle_prettifier","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"calculating-the-similarity-between-words-and","repo_url":"https://github.com/dyashkir/todo-ideas-links","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"calculating-the-similarity-between-words-and","repo_url":"https://github.com/nihitsaxena95/sentence-similarity-wordnet-sementic","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"calculating-the-similarity-between-words-and","repo_url":"https://github.com/zep283/Semantic_Similarity","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"semantic-similarity","task_name":"Semantic Similarity"},{"task_slug":"semantic-textual-similarity","task_name":"Semantic Textual Similarity"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"sentence-similarity","task_name":"Sentence Similarity"},{"task_slug":"word-similarity","task_name":"Word Similarity"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1802.05667","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}