Datasets › TURINGBENCH

TURINGBENCH

Introduced by Adaku Uchendu et al. in TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation27 Sep 2021 archive 2025-07-28

TuringBench is a benchmark environment that contains :

  • Benchmark tasks- Turing Test (i.e., human vs. machine) and Authorship Attribution: (i.e., who is the author of this texts?)
  • Datasets (Binary and Multi-class settings)
  • Website with leaderboard

The dataset has 20 labels (19 AI text-generators and human). We built this dataset by collecting 10K news articles (mostly Politics) from sources like CNN and only keeping articles with 200-400 words. Next, we used the Titles of these human-written articles to prompt the AI text-generators (ex: GPT-2, GROVER, etc.) to generate 10K articles each. This gives us a sum total of 200K articles and 20 labels. However, since there are two benchmark tasks - Turing Test and Authorship Attribution settings, we have all 20 labels in one dataset for the multi-class setting and only human vs. one AI text-generator, making 19 binary-class datasets.

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Binary text classification TURINGBENCH (Turing Test, GPT-3) GigaCheck (Mistral-7B) F1 score 0.9709 GigaCheck: Detecting LLM-generated Content — 2 Compare
Binary text classification TURINGBENCH (Turing Test, FAIR_wmt20) GigaCheck (Mistral-7B) F1 score 0.9966 GigaCheck: Detecting LLM-generated Content — 2 Compare

Papers archive 2025-07-28

2 shown of 2 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 18. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
GigaCheck: Detecting LLM-generated Content 0 2 31 Oct 2024 not harvested
TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation 3 2 27 Sep 2021 ran 0 of 7 samples (7 unverified)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

MIT

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • TURINGBENCH
  • TURINGBENCH (Turing Test, GPT-3)
  • TURINGBENCH (Turing Test, FAIR_wmt20)

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections