Browse State-of-the-Art › Sentence Similarity
Sentence Similarity
76 papers with code · 0 benchmarks · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 76 papers with code (194 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
14 Mar 2018 11 repositories listed Syntology ran 6 of 8 samples · 2 unverified · 6 pointer-only (licence)We introduce SentEval, a toolkit for evaluating the quality of universal sentence representations.
-
8 Apr 2020 4 repositories listed Syntology ran 4 of 8 samples · 4 unverifiedTransformer-based NLP models are trained using hundreds of millions or even billions of parameters, limiting their applicability in computationally constrained environments.
-
15 Feb 2018 4 repositories listedTo calculate the semantic similarity between words and sentences, the proposed method follows an edge-based approach using a lexical database.
-
26 Sep 2017 3 repositories listedWe propose a new generative model of sentences that first samples a prototype sentence from the training corpus and then edits it into a new sentence.
-
24 May 2023 2 repositories listed Syntology ran 2 of 6 samples · 4 unverifiedContrastive learning has been the dominant approach to train state-of-the-art sentence embeddings.
-
31 Jul 2020 2 repositories listedIn this paper, we challenge this assumption by showing that for domains with abundant unlabeled text, such as biomedicine, pretraining language models from scratch results in substantial gains over continual pretraining…
-
20 Dec 2018 2 repositories listedResearch on time-series similarity measures has emphasized the need for elastic methods which align the indices of pairs of time series and a plethora of non-parametric have been proposed for the task.
-
29 Aug 2018 2 repositories listedWe present a framework for building unsupervised representations of entities and their compositions, where each entity is viewed as a probability distribution rather than a vector embedding.
-
25 Jul 2017 2 repositories listedTo learn a semantic parser from denotations, a learning algorithm must search over a combinatorially large space of logical forms for ones consistent with the annotated denotations.
-
8 Nov 2016 2 repositories listedModeling the structure of coherent texts is a key NLP problem.
-
10 Dec 2024 1 repository listedHowever, the existence of such capability transfer between natural language and gene sequences/languages remains under explored.
-
26 Nov 2024 1 repository listedHere, we propose a novel Word pair-based Gaussian Sentence Similarity (WGSS) algorithm for calculating the semantic relation between two sentences.
-
21 Oct 2024 1 repository listedWe verify the efficacy of fine-tuning and conduct a series of experiments that assess the robustness of our method for low-resource scenarios.
-
17 Jul 2024 1 repository listedA challenge with word embedding is that as the vocabulary grows, the vector space's dimension increases, which can lead to a vast model size.
-
20 Jun 2024 1 repository listedWe assemble a broad Natural Language Understanding benchmark suite for the German language and consequently evaluate a wide array of existing German-capable models in order to create a better understanding of the…
-
4 Jun 2024 1 repository listedIn this work, we introduce OTTAWA, a novel Optimal Transport (OT)-based word aligner specifically designed to enhance the detection of hallucinations and omissions in MT systems.
-
4 Apr 2024 1 repository listedTo address this gap our paper introduces HunSum-2 an open-source Hungarian corpus suitable for training abstractive and extractive summarization models.
-
28 Jan 2024 1 repository listedWe apply a novel Mixture of Experts (MoE) extension pipeline to pretrained BERT models, where every multi-layer perceptron section is enlarged and copied into multiple distinct experts.
-
14 Jul 2023 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)We find that it can provide two aspects of information for the model, 1) it can help the model implicitly learn the location of semantic boundaries in continuous sign language videos, 2) it can help the model understand…
-
6 Jul 2023 1 repository listedWe show that this is also the case for sentence similarity, a fundamental task in multiple domains, e.
-
30 Jun 2023 1 repository listedMany self-supervised speech models (S3Ms) have been introduced over the last few years, improving performance and data efficiency on various speech tasks.
-
1 Jun 2023 1 repository listedWe investigate outlier dimensions and their relationship to anisotropy in multiple pre-trained multilingual language models.
-
24 May 2023 1 repository listed Syntology ran 2 of 4 samples · 2 unverified · 4 pointer-only (licence)Semantic textual similarity (STS), a cornerstone task in NLP, measures the degree of similarity between a pair of sentences, and has broad application in fields such as information retrieval and natural language…
-
17 May 2023 1 repository listed Syntology ran 0 of 2 samples · 2 unverifiedFurthermore, we propose a baseline model trained on this dataset, which outperforms model trained on the same data without images and BlenderBot.
-
24 Feb 2023 1 repository listedUnsupervised extractive summarization aims to extract salient sentences from a document as the summary without labeled data.
-
11 Jan 2023 1 repository listedPrior work in biomedical VLP has mostly relied on the alignment of single image and report pairs even though clinical notes commonly refer to prior images.
-
18 Dec 2022 1 repository listedIn this paper, we aim to help guide future designs of sentence representation learning methods by taking a closer look at contrastive SRL through the lens of isotropy, contextualization and learning dynamics.
-
21 Nov 2022 1 repository listedWe evaluate these models on real text classification datasets to show embeddings obtained from synthetic data training are generalizable to real datasets as well and thus represent an effective training strategy for…
-
24 Oct 2022 1 repository listedIn the field of natural language processing (NLP), continuous vector representations are crucial for capturing the semantic meanings of individual words.
-
29 Aug 2022 1 repository listedTo analyze this, we first train a classifier that identifies machine-written sentences, and observe that the linguistic features of the sentences identified as written by a machine are significantly different from those…
Syntology lines on 6 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections