Datasets › TURINGBENCH
TURINGBENCH
TuringBench is a benchmark environment that contains :
- Benchmark tasks- Turing Test (i.e., human vs. machine) and Authorship Attribution: (i.e., who is the author of this texts?)
- Datasets (Binary and Multi-class settings)
- Website with leaderboard
The dataset has 20 labels (19 AI text-generators and human). We built this dataset by collecting 10K news articles (mostly Politics) from sources like CNN and only keeping articles with 200-400 words. Next, we used the Titles of these human-written articles to prompt the AI text-generators (ex: GPT-2, GROVER, etc.) to generate 10K articles each. This gives us a sum total of 200K articles and 20 labels. However, since there are two benchmark tasks - Turing Test and Authorship Attribution settings, we have all 20 labels in one dataset for the multi-class setting and only human vs. one AI text-generator, making 19 binary-class datasets.
Benchmarks archive 2025-07-28
All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Binary text classification | TURINGBENCH (Turing Test, GPT-3) | GigaCheck (Mistral-7B) F1 score 0.9709 | GigaCheck: Detecting LLM-generated Content | — | 2 | Compare |
| Binary text classification | TURINGBENCH (Turing Test, FAIR_wmt20) | GigaCheck (Mistral-7B) F1 score 0.9966 | GigaCheck: Detecting LLM-generated Content | — | 2 | Compare |
Papers archive 2025-07-28
2 shown of 2 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 18. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| GigaCheck: Detecting LLM-generated Content | 0 | 2 | 31 Oct 2024 | not harvested |
| TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation | 3 | 2 | 27 Sep 2021 | ran 0 of 7 samples (7 unverified) |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- TURINGBENCH
- TURINGBENCH (Turing Test, GPT-3)
- TURINGBENCH (Turing Test, FAIR_wmt20)
3 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections