{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/to-word-senses-and-beyond-inducing-concepts","title":"To Word Senses and Beyond: Inducing Concepts with Contextualized Language Models","arxiv_id":"2406.20054","date":"2024-06-28","proceeding":null,"authors":["Bastien Liétard","Pascal Denis","Mikaella Keller"],"abstract":"Polysemy and synonymy are two crucial interrelated facets of lexical ambiguity. While both phenomena are widely documented in lexical resources and have been studied extensively in NLP, leading to dedicated systems, they are often being considered independently in practical problems. While many tasks dealing with polysemy (e.g. Word Sense Disambiguiation or Induction) highlight the role of word's senses, the study of synonymy is rooted in the study of concepts, i.e. meanings shared across the lexicon. In this paper, we introduce Concept Induction, the unsupervised task of learning a soft clustering among words that defines a set of concepts directly from data. This task generalizes Word Sense Induction. We propose a bi-level approach to Concept Induction that leverages both a local lemma-centric view and a global cross-lexicon view to induce concepts. We evaluate the obtained clustering on SemCor's annotated data and obtain good performance (BCubed F1 above 0.60). We find that the local and the global levels are mutually beneficial to induce concepts and also senses in our setting. Finally, we create static embeddings representing our induced concepts and use them on the Word-in-Context task, obtaining competitive performance with the State-of-the-Art.","url_abs":"https://arxiv.org/abs/2406.20054v2","url_pdf":"https://arxiv.org/pdf/2406.20054v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"lemma","task_name":"LEMMA"},{"task_slug":"word-sense-induction","task_name":"Word Sense Induction"}],"methods":[{"method_slug":"set","method_name":"SET"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2406.20054","atlas_url":"https://app.syntology.ai/?focus=2406.20054","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2406.20054"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/blietard/concept-induction","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":2,"unverified":2},"by_repo_kind":{"found_in_text":{"samples":4,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"452f936a6bdf7b80","entry":"get_algo_class","repo":"blietard/concept-induction","repo_kind":"found_in_text","path":"utilities/clustering.py","file_url":"https://github.com/blietard/concept-induction/blob/HEAD/utilities/clustering.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"452f936a6bdf7b80"}},{"code_sha256_prefix":"3e405281884794dc","entry":"verif_layers_arg","repo":"blietard/concept-induction","repo_kind":"found_in_text","path":"utilities/languagemodel.py","file_url":"https://github.com/blietard/concept-induction/blob/HEAD/utilities/languagemodel.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3e405281884794dc"}},{"code_sha256_prefix":"81b99eca69445465","entry":"get_tokenizer","repo":"blietard/concept-induction","repo_kind":"found_in_text","path":"utilities/languagemodel.py","file_url":"https://github.com/blietard/concept-induction/blob/HEAD/utilities/languagemodel.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"81b99eca69445465"}},{"code_sha256_prefix":"e29e021d7ad378d7","entry":"initialize_models","repo":"blietard/concept-induction","repo_kind":"found_in_text","path":"utilities/languagemodel.py","file_url":"https://github.com/blietard/concept-induction/blob/HEAD/utilities/languagemodel.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e29e021d7ad378d7"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}