Browse State-of-the-Art › Word Sense Disambiguation
Word Sense Disambiguation
150 papers with code · 16 benchmarks · 16 datasets archive 2025-07-28
The task of Word Sense Disambiguation (WSD) consists of associating words in context with their most suitable entry in a pre-defined sense inventory. The de-facto sense inventory for English in WSD is WordNet.. For example, given the word “mouse” and the following sentence:
“A mouse consists of an object held in one's hand, with one or more buttons.”
we would assign “mouse” with its electronic device sense (the 4th sense in the WordNet sense inventory).
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
16 leaderboard tables shown for this task, 16 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 16 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
16 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 150 papers with code (1,035 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
28 May 2020 67 repositories listed Syntology ran 15 of 65 samples · 50 unverified · 4 pointer-only (licence)By contrast, humans can generally perform a new language task from only a few examples or from simple instructions - something which current NLP systems still largely struggle to do.
-
23 Oct 2019 57 repositories listed Syntology ran 2 of 31 samples · 29 unverifiedTransfer learning, where a model is first pre-trained on a data-rich task before being fine-tuned on a downstream task, has emerged as a powerful technique in natural language processing (NLP).
-
5 Jun 2020 14 repositories listed Syntology ran 4 of 13 samples · 9 unverified · 3 pointer-only (licence)Recent progress in pre-trained neural language models has significantly improved the performance of many natural language processing (NLP) tasks.
-
11 Dec 2019 7 repositories listed Syntology ran 1 of 9 samples · 8 unverifiedLanguage models have become a key step to achieve state-of-the art results in many different Natural Language Processing (NLP) tasks.
-
14 Apr 2021 6 repositories listedThe approach significantly enhances the performance and interpretability of TM.
-
22 Aug 2016 4 repositories listed Syntology ran 4 of 4 samples · 0 unverified · 3 pointer-only (licence)We describe a transition-based parser for AMR that parses sentences left-to-right, in linear time.
-
28 Dec 2022 3 repositories listed Syntology ran 7 of 15 samples · 8 unverifiedFirst, we use synthetic language modeling tasks to understand the gap between SSMs and attention.
-
17 Feb 2022 3 repositories listed Syntology ran 5 of 5 samples · 0 unverified · 5 pointer-only (licence)But advancing the state-of-the-art across a broad set of natural language tasks has been hindered by training instabilities and uncertain quality during fine-tuning.
-
8 Dec 2021 3 repositories listedLanguage modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.
-
20 Aug 2019 3 repositories listedWord Sense Disambiguation (WSD) aims to find the exact sense of an ambiguous word in a particular context.
-
23 May 2023 2 repositories listedFurthermore, we show that instruction tuning with CoT Collection allows LMs to possess stronger few-shot learning capabilities on 4 domain-specific tasks, resulting in an improvement of +2.
-
7 Feb 2023 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Recently, Language Models (LMs) instruction-tuned on multiple tasks, also known as multitask-prompted fine-tuning (MT), have shown the capability to generalize to unseen tasks.
-
13 Jul 2022 2 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedTransformer models have recently emerged as one of the foundational models in natural language processing, and as a byproduct, there is significant recent interest and investment in scaling these models.
-
29 Jun 2022 2 repositories listedTo address this limitation, we propose SememeWSD Synonym (SWSDS) model to assign a different vector to every sense of polysemous words with the help of word sense disambiguation (WSD) and synonym set in OpenHowNet.
-
10 May 2022 2 repositories listed Syntology ran 0 of 16 samples · 16 unverifiedOur model also achieve strong results at in-context learning, outperforming 175B GPT-3 on zero-shot SuperGLUE and tripling the performance of T5-XXL on one-shot summarization.
-
29 Mar 2022 2 repositories listed Syntology ran 8 of 11 samples · 3 unverified · 4 pointer-only (licence)We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget.
-
15 Jun 2021 2 repositories listedWe present two supervised (pre-)training methods to incorporate gloss definitions from lexical resources into neural language models (LMs).
-
25 Apr 2021 2 repositories listedThe challenges with NLP systems with regards to tasks such as Machine Translation (MT), word sense disambiguation (WSD) and information retrieval make it imperative to have a labelled idioms dataset with classes such as…
-
18 Jan 2021 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedIn this work we study a mathematical formalization of this network motif and apply it to learning the correlational structure between words and their context in a corpus of unstructured text, a common natural language…
-
29 Oct 2020 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedIn this paper, we introduce an advanced Russian general language understanding evaluation benchmark -- RussianGLUE.
-
29 Apr 2020 2 repositories listedMeta-learning aims to solve this problem by training a model on a large number of few-shot tasks, with an objective to learn new tasks quickly from a small number of examples.
-
26 Nov 2019 2 repositories listedPre-trained word embeddings encode general word semantics and lexical regularities of natural language, and have proven useful across many NLP tasks, including word sense disambiguation, machine translation, and…
-
14 May 2019 2 repositories listedIn this article, we tackle the issue of the limited quantity of manually sense annotated corpora for the task of word sense disambiguation, by exploiting the semantic relationships between senses such as synonymy,…
-
19 Aug 2016 2 repositories listedIn this paper, we report a knowledge-based method for Word Sense Disambiguation in the domains of biomedical and clinical text.
-
7 Mar 2025 1 repository listed Syntology ran 0 of 2 samples · 2 unverifiedThe rise of generative chat-based Large Language Models (LLMs) over the past two years has spurred a race to develop systems that promise near-human conversational and reasoning experiences.
-
1 Mar 2025 1 repository listedLexical ambiguity is a major challenge in computational linguistic tasks, as limitations in proper sense identification lead to inefficient translation and question answering.
-
19 Dec 2024 1 repository listedThe model is based on Phi 2, an English-centric model of 2.
-
3 Jun 2024 1 repository listedThe fine-tuning of open-source large language models (LLMs) for machine translation has recently received considerable attention, marking a shift towards data-centric research from traditional neural machine translation.
-
11 May 2024 1 repository listedWe evaluate all existing models for contextualized Hebrew embeddings on a novel Hebrew homograph challenge sets that we deliver.
-
29 Apr 2024 1 repository listedExperimental results show that MMT models trained on our dataset exhibit a greater ability to exploit visual information than those trained on other MMT datasets.
Syntology lines on 14 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections