Home › Datasets › task › Semantic Similarity

Semantic Similarity datasets

archive 2025-07-28

17 datasets carry the task tag "Semantic Similarity" (the task itself: Semantic Similarity), ordered by the archive's paper count. Page 1 of 1: 17 shown of 17. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Semantic Similarity datasets 1–17 of 17

The MNIST database (Modified National Institute of Standards and Technology database) is a large collection of handwritten digits.
7,651 papers · 44 benchmarks
CoNLL-2003 is a named entity recognition dataset released as a part of CoNLL-2003 shared task: language-independent named entity recognition.
755 papers · 6 benchmarks
SICK (Sentences Involving Compositional Knowledge)
The Sentences Involving Compositional Knowledge (SICK) dataset is a dataset for compositional distributional semantics.
348 papers · 5 benchmarks
BIOSSES (Biomedical Semantic Similarity Estimation System)
The BIOSSES data set comprises total 100 sentence pairs all of which were selected from the "TAC2 Biomedical Summarization Track Training Data Set" .
38 papers · 2 benchmarks
Publicly available dataset of naturally occurring factual claims for the purpose of automatic claim verification.
21 papers · 0 benchmarks
CHIP-STS (Semantic Textual Similarity Dataset)
CHIP Semantic Textual Similarity, a dataset for sentence similarity in the non-i.i.d.
8 papers · 1 benchmark
The SUGARCREPE++ dataset evaluates the sensitivity of vision language models (VLMs) and unimodal language models (ULMs) to semantic and lexical alterations.
5 papers · 0 benchmarks
SV-Ident (Survey Variable Identification)
SV-Ident comprises 4,248 sentences from social science publications in English and German.
4 papers · 2 benchmarks
A benchmark dataset with 960 pairs of Chinese wOrd Similarity, where all the words have two morphemes in three Part of Speech (POS) tags with their human annotated similarity rather than relatedness.
3 papers · 0 benchmarks
Spoken versions of the Semantic Textual Similarity dataset for testing semantic sentence level embeddings.
3 papers · 0 benchmarks
This dataset contains information about Japanese word similarity including rare words.
2 papers · 0 benchmarks
LatamXIX (19th Century Latin American Spanish Newspaper Corpus with LLM OCR Correction)
A novel dataset of 19th-century Latin American press texts, which addresses the lack of specialized corpora for historical and linguistic analysis in this region.
2 papers · 0 benchmarks
Semantic Question Similarity in Arabic (NSURL-2019 Shared Task 8: Semantic Question Similarity in Arabic)
NSURL-2019 Shared Task 8: Semantic Question Similarity in Arabic This dataset contains 11,997 pairs of questions in MSA Arabic that are assigned either a label of 0, for no semantic similarity, or 1 otherwise.
2 papers · 0 benchmarks
Includes co-referent name string pairs along with their similarities.
1 paper · 0 benchmarks
Phrase in Context is a curated benchmark for phrase understanding and semantic search, consisting of three tasks of increasing difficulty: Phrase Similarity (PS), Phrase Retrieval (PR) and Phrase Sense Disambiguation (PSD).
1 paper · 0 benchmarks
Spanish Corpus XIX (19th Century Spanish Corpus)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Dataset Summary Speech Brown is a comprehensive, synthetic, and diverse paired speech-text dataset in 15 categories, covering a wide range of topics from fiction to religion.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.