{"url":"/dataset/turingbench","name":"TURINGBENCH","full_name":null,"description_markdown":"TuringBench is a benchmark environment that contains :\r\n\r\n- Benchmark tasks- Turing Test (i.e., human vs. machine) and Authorship Attribution: (i.e., who is the author of this texts?)\r\n- Datasets (Binary and Multi-class settings)\r\n- Website with leaderboard\r\n\r\nThe dataset has 20 labels (19 AI text-generators and human). We built this dataset by collecting 10K news articles (mostly Politics) from sources like CNN and only keeping articles with 200-400 words. Next, we used the Titles of these human-written articles to prompt the AI text-generators (ex: GPT-2, GROVER, etc.) to generate 10K articles each. This gives us a sum total of 200K articles and 20 labels. However, since there are two benchmark tasks - Turing Test and Authorship Attribution settings, we have all 20 labels in one dataset for the multi-class setting and only human vs. one AI text-generator, making 19 binary-class datasets.","description_withheld":null,"homepage":"https://turingbench.ist.psu.edu/","introduced_date":"2021-09-27","introduced_date_note":null,"introduced_by":{"paper":"/paper/turingbench-a-benchmark-environment-for","title":"TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation","first_author":"Adaku Uchendu","url":null},"license":{"name":"MIT","url":"https://opensource.org/license/mit"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Text Classification","url":"/task/text-classification","datasets_with_task":"/datasets/task/text-classification"},{"name":"Binary Classification","url":"/task/binary-classification","datasets_with_task":"/datasets/task/binary-classification"},{"name":"Binary text classification","url":"/task/binary-text-classification","datasets_with_task":"/datasets/task/binary-text-classification"},{"name":"Multi-class Classification","url":"/task/multi-class-classification","datasets_with_task":"/datasets/task/multi-class-classification"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["TURINGBENCH","TURINGBENCH (Turing Test, GPT-3)","TURINGBENCH (Turing Test, FAIR_wmt20)"],"data_loaders":[],"num_papers_in_archive":18,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/binary-text-classification-on-turingbench","task":"Binary text classification","dataset_variant":"TURINGBENCH (Turing Test, GPT-3)","rows":2,"metrics":["F1 score"],"first_row_in_archive_order":{"model":"GigaCheck (Mistral-7B)","paper":"/paper/gigacheck-detecting-llm-generated-content","metrics":{"F1 score":"0.9709"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/binary-text-classification-on-turingbench-1","task":"Binary text classification","dataset_variant":"TURINGBENCH (Turing Test, FAIR_wmt20)","rows":2,"metrics":["F1 score"],"first_row_in_archive_order":{"model":"GigaCheck (Mistral-7B)","paper":"/paper/gigacheck-detecting-llm-generated-content","metrics":{"F1 score":"0.9966"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/gigacheck-detecting-llm-generated-content","title":"GigaCheck: Detecting LLM-generated Content","date":"2024-10-31","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/turingbench-a-benchmark-environment-for","title":"TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation","date":"2021-09-27","rows_on_this_dataset":2,"code_links":3,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":7,"samples_ran":4,"samples_unverified":3,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":1,"samples_harvested":7,"samples_ran":4,"samples_unverified":3,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}