Browse State-of-the-Art › Word Similarity
Word Similarity
117 papers with code · 1 benchmark · 3 datasets archive 2025-07-28
Calculate a numerical score for the semantic similarity between two words.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| WS353 (3 rows) | Context-to-Vector | Using Context-to-Vector with Graph Retrofitting to Improve Word Embeddings | code | Syntology ran 1 of 2 samples · 1 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 117 papers with code (378 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
16 Jan 2013 84 repositories listed Syntology ran 19 of 62 samples · 43 unverified · 18 pointer-only (licence)We propose two novel model architectures for computing continuous vector representations of words from very large data sets.
-
15 Jul 2016 54 repositories listed Syntology ran 4 of 31 samples · 27 unverified · 4 pointer-only (licence)A vector representation is associated to each character n-gram; words being represented as the sum of these representations.
-
15 Feb 2018 4 repositories listedTo calculate the semantic similarity between words and sentences, the proposed method follows an edge-based approach using a lexical database.
-
7 Feb 2017 4 repositories listedMaybe the single most important goal of representation learning is making subsequent learning faster.
-
5 Feb 2017 4 repositories listed Syntology ran 2 of 26 samples · 24 unverifiedThe postprocessing is empirically validated on a variety of lexical-level intrinsic tasks (word similarity, concept categorization, word analogy) and sentence-level tasks (semantic textural similarity and { text…
-
30 Dec 2020 3 repositories listedIn this paper, we propose SemGloVe, which distills semantic co-occurrences from BERT into static GloVe word embeddings.
-
27 Aug 2018 3 repositories listedMultilingual Word Embeddings (MWEs) represent words from multiple languages in a single distributional vector space.
-
23 Mar 2018 3 repositories listedIn this paper, we propose a novel deep neural network architecture, Speech2Vec, for learning fixed-length vector representations of audio segments excised from a speech corpus, where the vectors contain semantic…
-
20 Oct 2023 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)But to achieve these results, LMs must be trained in distinctly un-human-like ways - requiring orders of magnitude more language data than children receive during development, and without perceptual or social context.
-
18 Jan 2021 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedIn this work we study a mathematical formalization of this network motif and apply it to learning the correlational structure between words and their context in a corpus of unstructured text, a common natural language…
-
17 Jan 2020 2 repositories listedThis paper presents a new technique for creating monolingual and cross-lingual meta-embeddings.
-
21 Oct 2018 2 repositories listedThis paper introduces the first dataset for evaluating English-Chinese Bilingual Contextual Word Similarity, namely BCWS (https://github.
-
18 Sep 2018 2 repositories listedContinuous word representation (aka word embedding) is a basic building block in many neural network-based models used in natural language processing tasks.
-
30 Aug 2018 2 repositories listedRecent work has demonstrated that embeddings of tree-like graphs in hyperbolic space surpass their Euclidean counterparts in performance by a large margin.
-
27 Aug 2018 2 repositories listedOur approach decouples learning the transformation from the source language to the target language into (a) learning rotations for language-specific embeddings to align them to a common space, and (b) learning a…
-
19 Jul 2018 2 repositories listedIn other words, we align words that are already determined to be related, along predefined concepts.
-
21 Apr 2018 2 repositories listedThe method consists of 3 steps as follows: (i) Expanding 1 or more dimension(s) on all the word vectors, filling with their representative value.
-
6 Jul 2017 2 repositories listedEvaluating these methods is also problematic, as rigorous quantitative evaluations in this space is limited, especially when compared with single-sense embeddings.
-
27 Apr 2017 2 repositories listedWord embeddings provide point representations of words containing useful semantic information.
-
11 Apr 2017 2 repositories listedThis paper describes Luminoso's participation in SemEval 2017 Task 2, "Multilingual and Cross-lingual Semantic Word Similarity", with a system based on ConceptNet.
-
17 Mar 2017 2 repositories listedAn evaluation of distributed word representation is generally conducted using a word similarity task and/or a word analogy task.
-
1 Dec 2016 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)Distributed representations of words have been shown to capture lexical semantics, as demonstrated by their effectiveness in word similarity and analogical relation tasks.
-
9 Jun 2015 2 repositories listedThen, based on this insight, we propose a novel framework WordRank that efficiently estimates word representations via robust ranking, in which the attention mechanism and robustness to noise are readily achieved via…
-
16 Jan 2025 1 repository listedTo address those issues, we propose a simple yet intuitive framework for how semantic shifts occur over multiple time periods by leveraging a similarity matrix between the embeddings of the same word through time.
-
13 May 2024 1 repository listedThe diffusion in the local graph manifold allows the exploration of the complex nonlinear geometry of word embeddings to capture word similarities based on paths of semantic association, over and above direct pairwise…
-
30 Mar 2024 1 repository listedTo this end, inspired by molecular phylogenetics, we propose a likelihood ratio test to determine if given languages are related based on the proportion of invariant character sites in the aligned wordlists applied…
-
20 Feb 2024 1 repository listedWhile BERT produces high-quality sentence embeddings, its pre-training computational cost is a significant drawback.
-
10 Jul 2023 1 repository listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)A promising candidate for the faithful embedding of data with varying structure is product manifolds of component spaces of different geometries (spherical, hyperbolic, or euclidean).
-
17 May 2023 1 repository listedThis similarity underestimation problem is particularly severe for highly frequent words.
-
19 Feb 2023 1 repository listedWe present a neural Sanskrit Natural Language Processing (NLP) toolkit named SanskritShala (a school of Sanskrit) to facilitate computational linguistic analyses for several tasks such as word segmentation,…
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections