Datasets › ChatGPT Paraphrases
ChatGPT Paraphrases
This is a dataset of paraphrases created by ChatGPT.
We used this prompt to generate paraphrases:
Generate 5 similar paraphrases for this question, show it like a numbered list without commentaries: {text}
This dataset is based on the Quora paraphrase question, texts from the SQUAD 2.0 and the CNN news dataset.
We generated 5 paraphrases for each sample, totally this dataset has about 350k data rows. You can make 30 rows from a row from each sample. In this way you can make 10.5 millions train pairs (350k rows with 5 paraphrases -> 6x5x350000 = 10.5 millions of bidirected or 6x5x350000/2 = 5.25 millions of unique pairs).
We used:
-
231927 questions from the Quora dataset
-
92005 texts from the Squad 2.0 dataset
-
29110 texts from the CNN news dataset
Structure of the dataset:
-
text column - an original sentence or question from the datasets
-
paraphrases - a list of 5 paraphrases
-
category - question / sentence
-
source - quora / squad_2 / cnn_news
Benchmarks archive 2025-07-28
No leaderboard in the archive resolves to this dataset.
Papers archive 2025-07-28
No paper in the archive has a leaderboard row on this dataset.
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
openrail
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- ChatGPT Paraphrases
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections