{"url":"/dataset/identifying-machine-paraphrased-plagiarism","name":"Machine Prarphrase Corpus (MPC)","full_name":null,"description_markdown":"This dataset is used to train and evaluate models for the detection of machine-paraphrased text.\r\n\r\nThe training set consists of 200,767 paragraphs (98,282 original, 102,485 paraphrased) extracted from 8,024 Wikipedia (English) articles (4,012 original, 4,012 paraphrased using the SpinBot API).\r\n\r\nThe test set is divided into 3 subsets: one created from preprints of research papers on arXiv, one from graduation theses, and one from Wikipedia articles. Additionally, different marchine-paraphrasing methods were used.\r\n\r\nTest sets:\r\n```\r\nSpinBot: \r\n    arXiv         - Original - 20,966;    Spun - 20,867\r\n    Theses        - Original - 5,226;        Spun - 3,463\r\n    Wikipedia    - Original - 39,241;    Spun - 40,729\r\n    \r\nSpinnerChief-4W: \r\n    arXiv         - Original - 20,966;    Spun - 21,671\r\n    Theses        - Original - 2,379;        Spun - 2,941\r\n    Wikipedia    - Original - 39,241;    Spun - 39,618\r\n    \r\nSpinnerChief-2W: \r\n    arXiv         - Original - 20,966;    Spun - 21,719\r\n    Theses        - Original - 2,379;        Spun - 2,941\r\n    Wikipedia    - Original - 39,241;    Spun - 39,697\r\n```","description_withheld":null,"homepage":"https://doi.org/10.5281/zenodo.3608000","introduced_date":"2021-03-22","introduced_date_note":null,"introduced_by":{"paper":"/paper/identifying-machine-paraphrased-plagiarism","title":"Identifying Machine-Paraphrased Plagiarism","first_author":"Jan Philip Wahle","url":null},"license":{"name":"Attribution 4.0 International","url":"https://creativecommons.org/licenses/by/4.0/legalcode"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Paraphrase Identification","url":"/task/paraphrase-identification","datasets_with_task":"/datasets/task/paraphrase-identification"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Machine Prarphrase Corpus (MPC)"],"data_loaders":[{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/jpwahle/machine-paraphrase-dataset","frameworks":["tf","pytorch","jax"]}],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}