{"url":"/dataset/chatgpt-paraphrases","name":"ChatGPT Paraphrases","full_name":null,"description_markdown":"This is a dataset of paraphrases created by ChatGPT.\r\n\r\n**We used this prompt to generate paraphrases:**                         \r\nGenerate 5 similar paraphrases for this question, show it like a numbered list without commentaries: *{text}*\r\n\r\nThis dataset is based on the [Quora paraphrase question](https://www.kaggle.com/competitions/quora-question-pairs), texts from the [SQUAD 2.0](https://huggingface.co/datasets/squad_v2) and the [CNN news dataset](https://huggingface.co/datasets/cnn_dailymail).\r\n\r\nWe generated 5 paraphrases for each sample, totally this dataset has about 350k data rows. You can make 30 rows from a row \r\nfrom each sample. In this way you can make 10.5 millions train pairs (350k rows with 5 paraphrases -> 6x5x350000 = 10.5 millions of bidirected or 6x5x350000/2 = 5.25 millions of unique pairs).\r\n\r\n**We used:**\r\n\r\n- 231927 questions from the Quora dataset\r\n\r\n- 92005 texts from the Squad 2.0 dataset\r\n\r\n- 29110 texts from the CNN news dataset\r\n\r\n**Structure of the dataset:**\r\n\r\n- text column - an original sentence or question from the datasets\r\n\r\n- paraphrases - a list of 5 paraphrases\r\n\r\n- category - question / sentence\r\n\r\n- source - quora / squad_2 / cnn_news","description_withheld":null,"homepage":"https://huggingface.co/datasets/humarin/chatgpt-paraphrases","introduced_date":"2023-03-15","introduced_date_note":null,"introduced_by":null,"license":{"name":"openrail","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Paraphrase Generation","url":"/task/paraphrase-generation","datasets_with_task":"/datasets/task/paraphrase-generation"},{"name":"text2text-generation","url":"/task/text2text-generation","datasets_with_task":"/datasets/task/text2text-generation"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["ChatGPT Paraphrases"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}