Browse State-of-the-Art › STS
STS
129 papers with code · 0 benchmarks · 4 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 129 papers with code (334 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
27 Aug 2019 64 repositories listed Syntology ran 20 of 58 samples · 38 unverified · 11 pointer-only (licence)However, it requires that both sentences are fed into the network, which causes a massive computational overhead: Finding the most similar pair in a collection of 10, 000 sentences requires about 50 million inference…
-
18 Apr 2021 23 repositories listed Syntology ran 17 of 30 samples · 13 unverified · 19 pointer-only (licence)This paper presents SimCSE, a simple contrastive learning framework that greatly advances state-of-the-art sentence embeddings.
-
14 Apr 2021 6 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedLearning sentence embeddings often requires a large amount of labeled data.
-
13 Oct 2022 5 repositories listed Syntology ran 3 of 13 samples · 10 unverifiedMTEB spans 8 embedding tasks covering a total of 58 datasets and 112 languages.
-
28 Aug 2018 5 repositories listedA subset of MedSTS (MedSTS_ann) containing 1, 068 sentence pairs was annotated by two medical experts with semantic similarity scores of 0-5 (low to high similarity).
-
7 Apr 2020 3 repositories listedAlthough several benchmark datasets for those tasks have been released in English and a few other languages, there are no publicly available NLI or STS datasets in the Korean language.
-
SemEval-2017 Task 1: Semantic Textual Similarity - Multilingual and Cross-lingual Focused Evaluation31 Jul 2017 3 repositories listedSemantic Textual Similarity (STS) measures the meaning similarity of sentences.
-
14 Jun 2024 2 repositories listedSemantic Textual Similarity (STS) constitutes a critical research direction in computational linguistics and serves as a key indicator of the encoding capabilities of embedding models.
-
8 Jun 2024 2 repositories listedSince the introduction of BERT and RoBERTa, research on Semantic Textual Similarity (STS) has made groundbreaking progress.
-
13 Feb 2024 2 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)To our knowledge, this is the first representation learning method devoid of traditional language models for understanding sentence and document semantics, marking a stride closer to human-like textual comprehension.
-
9 Nov 2023 2 repositories listed Syntology ran 7 of 7 samples · 0 unverifiedMost recent studies employed large language models (LLMs) to learn sentence embeddings.
-
22 Sep 2023 2 repositories listedThis novel approach effectively mitigates the adverse effects of the saturation zone in the cosine function, which can impede gradient and hinder optimization processes.
-
8 Oct 2022 2 repositories listedContrastive learning has been extensively studied in sentence embedding learning, which assumes that the embeddings of different views of the same sentence are closer.
-
ESimCSE: Enhanced Sample Building Method for Contrastive Learning of Unsupervised Sentence Embedding9 Sep 2021 2 repositories listedUnsup-SimCSE takes dropout as a minimal data augmentation method, and passes the same input sentence to a pre-trained Transformer encoder (with dropout turned on) twice to obtain the two corresponding embeddings to…
-
9 Sep 2021 2 repositories listedContrastive learning has been gradually applied to learn high-quality unsupervised sentence embedding.
-
19 Aug 2021 2 repositories listedTo support our investigation, we establish a new sentence representation transfer benchmark, SentGLUE, which extends the SentEval toolkit to nine tasks from the GLUE benchmark.
-
22 Mar 2021 2 repositories listedThis study provides an efficient approach for using text data to calculate patent-to-patent (p2p) technological similarity, and presents a hybrid framework for leveraging the resulting p2p similarity for applications…
-
27 Nov 2020 2 repositories listedIn this paper, we propose FFCI, a framework for fine-grained summarization evaluation that comprises four elements: faithfulness (degree of factual consistency with the source), focus (precision of summary content…
-
30 Apr 2019 2 repositories listedRecent literature suggests that averaged word vectors followed by simple post-processing outperform many deep learning methods on semantic textual similarity tasks.
-
19 May 2025 1 repository listedPrevious studies usually focus on prompt engineering to guide LLMs to encode the core semantic information of the sentence into the embedding of the last token.
-
4 May 2025 1 repository listedEmbedding layers in transformer-based NLP models typically account for the largest share of model parameters, scaling with vocabulary size but not yielding performance gains proportional to scale.
-
12 Mar 2025 1 repository listedSuch sentence embeddings can be further enhanced by domain adaptation that adapts a backbone model to a specific domain.
-
9 Mar 2025 1 repository listedVisual Place Recognition (VPR) is a crucial capability for long-term autonomous robots, enabling them to identify previously visited locations using visual information.
-
16 Feb 2025 1 repository listedEmbedding models play a crucial role in representing and retrieving information across various NLP applications.
-
26 Nov 2024 1 repository listed Syntology ran 1 of 2 samples · 1 unverifiedThompson sampling (TS) has optimal regret and excellent empirical performance in multi-armed bandit problems.
-
26 Nov 2024 1 repository listedIn this reproducibility study, we implement and evaluate both versions of 2D Matryoshka Training on STS tasks and extend our analysis to retrieval tasks.
-
18 Oct 2024 1 repository listed Syntology ran 1 of 3 samples · 2 unverified · 3 pointer-only (licence)Training-free embedding methods directly leverage pretrained large language models (LLMs) to embed text, bypassing the costly and complex procedure of contrastive learning.
-
13 Sep 2024 1 repository listedThe results on the monolingual tasks confirmed that our representations exhibited a competitive performance compared to that of the previous study for the context-aware lexical semantic tasks and outperformed it for STS…
-
28 Aug 2024 1 repository listedThis paper examines the Code-Switching (CS) phenomenon where two languages intertwine within a single utterance.
-
1 Aug 2024 1 repository listedWhile Large Language Models show remarkable performance in natural language understanding, their resource-intensive nature makes them less accessible.
Syntology lines on 8 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections