Browse State-of-the-Art › Word Sense Induction
Word Sense Induction
19 papers with code · 1 benchmark · 1 dataset archive 2025-07-28
Word sense induction (WSI) is widely known as the “unsupervised version” of WSD. The problem states as: Given a target word (e.g., “cold”) and a collection of sentences (e.g., “I caught a cold”, “The weather is cold”) that use the word, cluster the sentences according to their different senses/meanings. We do not need to know the sense/meaning of each cluster, but sentences inside a cluster should have used the target words with the same sense.
Description from NLP Progress
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| SemEval 2010 WSI (5 rows) | BERT+DP | Towards better substitution-based word sense induction | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
19 shown of 19 papers with code (107 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
25 Feb 2015 3 repositories listedRecently proposed Skip-gram model is a powerful method for learning high-dimensional word representations that capture rich semantic relationships between words.
-
28 Sep 2022 2 repositories listedWe present RuDSI, a new benchmark for word sense induction (WSI) in Russian.
-
29 May 2019 2 repositories listedWord sense induction (WSI) is the task of unsupervised clustering of word usages within a sentence to distinguish senses.
-
6 Jul 2017 2 repositories listedEvaluating these methods is also problematic, as rigorous quantitative evaluations in this space is limited, especially when compared with single-sense embeddings.
-
19 Feb 2024 1 repository listedOur evaluation is performed across different languages on eight available benchmarks for LSC, and shows that (i) APD outperforms other approaches for GCD; (ii) XL-LEXEME outperforms other contextualized models for WiC,…
-
19 Dec 2022 1 repository listedScholarly text is often laden with jargon, or specialized language that can facilitate efficient in-group communication within fields but hinder understanding for out-groups.
-
7 Jun 2022 1 repository listedLexical substitution, i.
-
25 Jan 2021 1 repository listedTo avoid the "meaning conflation deficiency" of word embeddings, a number of models have aimed to embed individual word senses.
-
1 Nov 2020 1 repository listedThe task of Diachronic Word Sense Induction (DWSI) aims to identify the meaning of words from their context, taking the temporal dimension into account.
-
1 Nov 2020 1 repository listedThe recent paradigm shift to contextual word embeddings has seen tremendous success across a wide range of down-stream tasks.
-
18 Oct 2020 1 repository listedMuch as the social landscape in which languages are spoken shifts, language too evolves to suit the needs of its users.
-
22 Nov 2018 1 repository listedThus, we aim to eliminate these requirements and solve the sense granularity problem by proposing AutoSense, a latent variable model based on two observations: (1) senses are represented as a distribution over topics,…
-
26 Aug 2018 1 repository listedAn established method for Word Sense Induction (WSI) uses a language model to predict probable substitutes for target words, and induces senses by clustering these resulting substitute vectors.
-
6 May 2018 1 repository listedThe paper reports our participation in the shared task on word sense induction and disambiguation for the Russian language (RUSSE-2018).
-
1 Jul 2017 1 repository listedThe key idea is to utilize word sememes to capture exact meanings of a word within specific contexts accurately.
-
24 Apr 2017 1 repository listedThis paper presents a new graph-based approach that induces synsets using synonymy dictionaries and word embeddings.
-
1 Apr 2017 1 repository listedTo evaluate our method we construct two 600-word testsets for word-to-synset matching in French and Russian using native speakers and evaluate the performance of our method along with several other recent approaches.
-
1 Jun 2013 1 repository listed
-
1 Jul 2012 1 repository listed
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections