Browse State-of-the-Art › text similarity
text similarity
100 papers with code · 0 benchmarks · 3 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 100 papers with code (271 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
21 Mar 2018 6 repositories listed Syntology ran 7 of 16 samples · 9 unverified · 1 pointer-only (licence)Prior work either simply aggregates the similarity of all possible pairs of regions and words without attending differentially to more and less important words or regions, or uses a multi-step attentional process to…
-
28 Nov 2023 3 repositories listedThis paper introduces RETSim (Resilient and Efficient Text Similarity), a lightweight, multilingual deep learning model trained to produce robust metric embeddings for near-duplicate text retrieval, clustering, and…
-
21 Mar 2023 3 repositories listed Syntology ran 5 of 6 samples · 1 unverifiedOpen-vocabulary semantic segmentation presents the challenge of labeling each pixel within an image based on a wide range of text descriptions.
-
28 Jun 2025 2 repositories listedWe introduce ActAlign, a zero-shot framework that formulates video classification as sequence alignment.
-
31 Mar 2024 2 repositories listedIn this work, we introduce Self-Contrast, a feedback-free large language model alignment method via exploiting extensive self-generated negatives.
-
19 Apr 2023 2 repositories listedWe also propose a scale-dot product attention mechanism to capture the similarity between title features and textual features.
-
8 Oct 2022 2 repositories listedContrastive learning has been extensively studied in sentence embedding learning, which assumes that the embeddings of different views of the same sentence are closer.
-
ESimCSE: Enhanced Sample Building Method for Contrastive Learning of Unsupervised Sentence Embedding9 Sep 2021 2 repositories listedUnsup-SimCSE takes dropout as a minimal data augmentation method, and passes the same input sentence to a pre-trained Transformer encoder (with dropout turned on) twice to obtain the two corresponding embeddings to…
-
9 Sep 2021 2 repositories listedContrastive learning has been gradually applied to learn high-quality unsupervised sentence embedding.
-
1 Nov 2020 2 repositories listedObtaining such a corpus from crowdworkers, however, has been shown to be ineffective since (i) workers usually lack domain-specific expertise to conduct the task with sufficient quality, and (ii) the standard approach…
-
8 Feb 2020 2 repositories listedThis paper proposes a chatbot framework that adopts a hybrid model which consists of a knowledge graph and a text similarity model.
-
12 Aug 2019 2 repositories listedWe propose a novel framework that achieves remarkable matching performance with acceptable model complexity.
-
15 Sep 2017 2 repositories listedThis network is composed of compare mechanism, two-staged CNN architecture with attention mechanism, and a prediction layer.
-
11 Jun 2025 1 repository listedDual encoder Vision-Language Models (VLM) such as CLIP are widely used for image-text retrieval tasks.
-
16 May 2025 1 repository listedEditing images using natural language instructions has become a natural and expressive way to modify visual content; yet, evaluating the performance of such models remains challenging.
-
5 May 2025 1 repository listedCombining the above two motivations, we propose a new \textbf{J}oint \textbf{T}ensor representation modulus constraint and \textbf{C}ross-attention unsupervised contrastive learning \textbf{S}entence \textbf{E}mbedding…
-
14 Apr 2025 1 repository listedLiterature review tables are essential for summarizing and comparing collections of scientific papers.
-
25 Mar 2025 1 repository listedRecent advancements in AI-driven conversational agents have exhibited immense potential of AI applications.
-
9 Mar 2025 1 repository listedVisual Place Recognition (VPR) is a crucial capability for long-term autonomous robots, enabling them to identify previously visited locations using visual information.
-
17 Jan 2025 1 repository listedHowever, their handcrafted generic descriptions fail to capture the diverse range of anomalies that may emerge in different objects, and simple patch-level image-text matching often struggles to localize anomalous…
-
16 Dec 2024 1 repository listedWe introduce Speech Information Retrieval (SIR), a new long-context task for Speech Large Language Models (Speech LLMs), and present SPIRAL, a 1, 012-sample benchmark testing models' ability to extract critical details…
-
9 Dec 2024 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedIn the peer review process of top-tier machine learning (ML) and artificial intelligence (AI) conferences, reviewers are assigned to papers through automated methods.
-
26 Nov 2024 1 repository listedIn this reproducibility study, we implement and evaluate both versions of 2D Matryoshka Training on STS tasks and extend our analysis to retrieval tasks.
-
31 Oct 2024 1 repository listedCross-modal text-molecule retrieval model aims to learn a shared feature space of the text and molecule modalities for accurate similarity calculation, which facilitates the rapid screening of molecules with specific…
-
18 Oct 2024 1 repository listed Syntology ran 1 of 3 samples · 2 unverified · 3 pointer-only (licence)Training-free embedding methods directly leverage pretrained large language models (LLMs) to embed text, bypassing the costly and complex procedure of contrastive learning.
-
17 Oct 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverifiedEffective approaches that can scale embedding model depth (i.
-
12 Oct 2024 1 repository listedRecent progress in audio pre-trained models and large language models (LLMs) has significantly enhanced audio understanding and textual reasoning capabilities, making improvements in AAC possible.
-
4 Sep 2024 1 repository listedWhile these models, employing different pooling and attention strategies, have achieved state-of-the-art performance on public embedding benchmarks, questions still arise about what constitutes an effective design for…
-
25 Jul 2024 1 repository listedStarting from the objective of positive reframing, we first design positive sentiment reward and content preservation reward to encourage the model to transform the negative expressions of the original text while…
-
20 Jul 2024 1 repository listedAnother limitation of multilingual sentence encoders is the trade-off between monolingual and cross-lingual performance.
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections