Home › Datasets › task › Paraphrase Identification
Paraphrase Identification datasets
archive 2025-07-28
18 datasets carry the task tag "Paraphrase Identification" (the task itself: Paraphrase Identification), ordered by the archive's paper count. Page 1 of 1: 18 shown of 18. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Paraphrase Identification datasets 1–18 of 18
The IMDb Movie Reviews dataset is a binary sentiment analysis dataset consisting of 50,000 reviews from the Internet Movie Database (IMDb) labeled as positive or negative.
1,787 papers · 9 benchmarks
PAWS-X contains 23,659 human translated PAWS evaluation pairs and 296,406 machine translated training pairs in six typologically distinct languages: French, Spanish, German, Chinese, Japanese, and Korean.
172 papers · 0 benchmarks
PAWS (Paraphrase Adversaries from Word Scrambling)
Paraphrase Adversaries from Word Scrambling (PAWS) is a dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of…
159 papers · 0 benchmarks
The Yelp Dataset is a valuable resource for academic research, teaching, and learning.
86 papers · 15 benchmarks
WikiHop is a multi-hop question-answering dataset.
66 papers · 2 benchmarks
Quora Question Pairs (QQP) dataset consists of over 400,000 question pairs, and each question pair is annotated with a binary value indicating whether the two questions are paraphrase of each other.
55 papers · 8 benchmarks
PIT (Paraphrase and Semantic Similarity in Twitter)
Paraphrase and Semantic Similarity in Twitter (PIT) presents a constructed Twitter Paraphrase Corpus that contains 18,762 sentence pairs.
22 papers · 1 benchmark
Paralex learns from a collection of 18 million question-paraphrase pairs scraped from WikiAnswers.
20 papers · 1 benchmark
FLUE (French Language Understanding Evaluation)
FLUE is a French Language Understanding Evaluation benchmark.
12 papers · 0 benchmarks
TURL (Twitter News URL Corpus)
Twitter News URL Corpus is a human-labeled paraphrase corpus to date of 51,524 sentence pairs and the first cross-domain benchmarking for automatic paraphrase identification.
6 papers · 1 benchmark
Finnish Paraphrase Corpus is a fully manually annotated paraphrase corpus for Finnish containing 53,572 paraphrase pairs harvested from alternative subtitles and news headings.
4 papers · 0 benchmarks
SV-Ident (Survey Variable Identification)
SV-Ident comprises 4,248 sentences from social science publications in English and German.
4 papers · 2 benchmarks
PARADE contains paraphrases that overlap very little at the lexical and syntactic level but are semantically equivalent based on computer science domain knowledge, as well as non-paraphrases that overlap greatly at the lexical and…
3 papers · 0 benchmarks
AP (Adversarial Paraphrase)
This is a paraphrasing dataset created using the adversarial paradigm.
1 paper · 1 benchmark
This is a benchmark for neural paraphrase detection, to differentiate between original and machine-generated content.
1 paper · 0 benchmarks
For more details see https://huggingface.co/datasets/jpwahle/autoregressive-paraphrase-dataset
1 paper · 0 benchmarks
This dataset is used to train and evaluate models for the detection of machine-paraphrased text.
1 paper · 0 benchmarks
Translated SNLI Dataset in Marathi A translated version of the SNLI dataset in Marathi, designed for Semantic Textual Similarity (STS) tasks.
1 paper · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.