Home › Datasets › task › Word Sense Disambiguation
Word Sense Disambiguation datasets
archive 2025-07-28
16 datasets carry the task tag "Word Sense Disambiguation" (the task itself: Word Sense Disambiguation), ordered by the archive's paper count. Page 1 of 1: 16 shown of 16. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Word Sense Disambiguation datasets 1–16 of 16
BIG-bench (Beyond the Imitation Game Benchmark)
The Beyond the Imitation Game Benchmark (BIG-bench) is a collaborative benchmark intended to probe large language models and extrapolate their future capabilities.
349 papers · 121 benchmarks
WiC is a benchmark for the evaluation of context-sensitive word embeddings.
206 papers · 3 benchmarks
The Evaluation framework of Raganato et al.
109 papers · 3 benchmarks
There are now many computer programs for automatically determining the sense of a word in context (Word Sense Disambiguation or WSD).
19 papers · 0 benchmarks
FLUE (French Language Understanding Evaluation)
FLUE is a French Language Understanding Evaluation benchmark.
12 papers · 0 benchmarks
WiC-TSV (Words-in-Context: Target Sense Verification)
WiC-TSV is a new multi-domain evaluation benchmark for Word Sense Disambiguation.
9 papers · 2 benchmarks
RUSSE (Russian Words in Context (based on RUSSE))
WiC: The Word-in-Context Dataset A reliable benchmark for the evaluation of context-sensitive word embeddings.
8 papers · 1 benchmark
SubjQA is a question answering dataset that focuses on subjective (as opposed to factual) questions and answers.
8 papers · 0 benchmarks
The CoarseWSD-20 dataset is a coarse-grained sense disambiguation dataset built from Wikipedia (nouns only) targeting 2 to 5 senses of 20 ambiguous words.
6 papers · 0 benchmarks
FEWS (FEWS: Large-Scale, Low-Shot Word Sense Disambiguation with the Dictionary)
FEWS (Few-shot Examples of Word Senses) is a few-shot dataset for English Word Sense Disambiguation (WSD) gathered from Wiktionary, an online, crowd-sourced dictionary.
4 papers · 1 benchmark
Verse is a new dataset that augments existing multimodal datasets (COCO and TUHOI) with sense labels.
4 papers · 0 benchmarks
Intended to provide freely available data sets in various formats together with basic annotation to be useful for applications in computational linguistics, translation studies and cross-linguistic corpus studies.
3 papers · 0 benchmarks
The Parallel Meaning Bank (PMB), developed at the University of Groningen and building upon the Groningen Meaning Bank, comprises sentences and texts in raw and tokenised format, syntactic analysis, word senses, thematic roles, reference…
2 papers · 0 benchmarks
ReviewQA is a question-answering dataset based on hotel reviews.
2 papers · 0 benchmarks
The Part-Whole Relations dataset is a dataset of semantic relations between entities.
1 paper · 0 benchmarks
SBU-WSD-Corpus is a corpus for Persian Word Sense Disambiguation (WSD).
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.