Browse State-of-the-Art › Semantic Retrieval
Semantic Retrieval
26 papers with code · 1 benchmark · 3 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Contract Discovery (6 rows) | Human baseline | Contract Discovery: Dataset and a Few-Shot Semantic Retrieval... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
26 shown of 26 papers with code (86 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
24 Aug 2023 2 repositories listedIn this work, we use a multilingual knowledge distillation approach to train BERT models to produce sentence embeddings for Ancient Greek text.
-
2 Jul 2023 2 repositories listedIn response, we introduce MedCPT, a first-of-its-kind Contrastively Pre-trained Transformer model for zero-shot semantic IR in biomedicine.
-
17 Sep 2019 2 repositories listed Syntology ran 6 of 17 samples · 11 unverifiedIn this work, we give general guidelines on system design for MRS by proposing a simple yet effective pipeline system with special consideration on hierarchical semantic retrieval at both paragraph and sentence level,…
-
21 May 2025 1 repository listedLarge Language Models (LLMs) have demonstrated their potential in hardware design tasks, such as Hardware Description Language (HDL) generation and debugging.
-
19 Feb 2025 1 repository listedIn this demonstration, we present AnDB, an AI-native database that supports traditional OLTP workloads and innovative AI-driven tasks, enabling unified semantic analysis across structured and unstructured data.
-
13 Jan 2025 1 repository listedSemantic retrieval (also known as dense retrieval) based on textual data has been extensively studied for both web search and product search application fields, where the relevance of a query and a potential target…
-
13 Dec 2024 1 repository listedTo address these challenges, optimizations for long-context inference have been developed, centered around the KV cache.
-
23 Jul 2024 1 repository listedBy incorporating emotional support strategies, we aim to enrich the model's capabilities in both cognitive and affective empathy, leading to a more nuanced and comprehensive empathetic response.
-
17 Jul 2024 1 repository listedMost existing Low-light Image Enhancement (LLIE) methods either directly map Low-Light (LL) to Normal-Light (NL) images or use semantic or illumination maps as guides.
-
11 Jun 2024 1 repository listed Syntology ran 9 of 10 samples · 1 unverifiedWords have been represented in a high-dimensional vector space that encodes their semantic similarities, enabling downstream applications such as retrieving synonyms, antonyms, and relevant contexts.
-
21 Feb 2024 1 repository listedIn the last years' digitalization process, the creation and management of documents in various domains, particularly in Public Administration (PA), have become increasingly complex and diverse.
-
30 Oct 2023 1 repository listed Syntology ran 4 of 4 samples · 0 unverifiedIn this paper, we propose M4LE, a Multi-ability, Multi-range, Multi-task, Multi-domain benchmark for Long-context Evaluation.
-
16 Oct 2023 1 repository listedWe demonstrate that LLMs semantic retrieval and reasoning abilities on problem-specific tasks can be applied to large textual archives that have not been part of the its training data.
-
25 May 2023 1 repository listedInspired by this, we replace the semantic retrieval in Retro with a surface-level method based on BM25, obtaining a significant reduction in perplexity.
-
29 Nov 2022 1 repository listedThey have overlooked the wide characteristic changes of different classes and can not model abundant intra-class variations for generations.
-
16 Oct 2022 1 repository listed Syntology ran 0 of 3 samples · 3 unverifiedThough offering amazing contextualized token-level representations, current pre-trained language models take less attention on accurately acquiring sentence-level representation during their self-supervised pre-training.
-
26 Jul 2022 1 repository listedOur sentence encoder can be trained in less than a day on a single graphics card, achieving high performance on a diverse set of sentence-level tasks.
-
19 Jul 2022 1 repository listedWhile contextualized word embeddings have been a de-facto standard, learning contextualized phrase embeddings is less explored and being hindered by the lack of a human-annotated benchmark that tests machine…
-
28 May 2022 1 repository listedIn this work, we will show that the inferior standard of accuracy draws from human annotations (leave-one-out) are not appropriate for machine-generated captions.
-
28 Apr 2022 1 repository listedThus, we model semantics by means of hidden random variables and define the semantic communication task as the data-reduced and reliable transmission of messages over a communication channel such that semantics is best…
-
15 Mar 2022 1 repository listedHow to learn highly compact yet effective sentence representation?
-
5 Jul 2021 1 repository listedWe compare the alignment performance using our proposed evaluation metrics to the semantic retrieval task commonly used to evaluate VGS models.
-
8 Mar 2021 1 repository listedWe believe it is the right time to survey current status, learn from existing methods, and gain some insights for future development.
-
22 Dec 2020 1 repository listedThis layer is shown to minimize a penalized term of the Wasserstein distance between the learned continuous image features and the optimal half-half bit distribution.
-
10 Nov 2019 1 repository listedWe propose a new shared task of semantic retrieval from legal texts, in which a so-called contract discovery is to be performed, where legal clauses are extracted from documents, given a few examples of similar clauses…
-
15 Apr 2019 1 repository listed Syntology ran 0 of 3 samples · 3 unverifiedA number of recent studies have started to investigate how speech systems can be trained on untranscribed speech by leveraging accompanying images at training time.
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections