{"url":"/dataset/roft","name":"RoFT","full_name":"Real or Fake Text","description_markdown":"RoFT is a dataset of 21,000 human annotations of generated text. The task is \"Boundary detection\" i.e. given a passage that starts off as human written, determine when the text transitions to being machine generated. The dataset also includes error annotations using the taxonomy introduced in the paper. The data can be used to train automatic detection systems, train automatic error correction, analyze visibility of model errors, and compare performance across models. Data was collected using http://roft.io. \r\n\r\nModels: GPT2, GPT2-XL, CTRL, GPT3 \"Davinci\"\r\n\r\nGenres: News, Stories, Recipes, Speeches","description_withheld":null,"homepage":"https://github.com/liamdugan/human-detection","introduced_date":"2022-12-24","introduced_date_note":null,"introduced_by":{"paper":"/paper/real-or-fake-text-investigating-human-ability","title":"Real or Fake Text?: Investigating Human Ability to Detect Boundaries Between Human-Written and Machine-Generated Text","first_author":"Liam Dugan","url":null},"license":{"name":"MIT","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Boundary Detection","url":"/task/boundary-detection","datasets_with_task":"/datasets/task/boundary-detection"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["RoFT"],"data_loaders":[],"num_papers_in_archive":3,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/boundary-detection-on-roft","task":"Boundary Detection","dataset_variant":"RoFT","rows":4,"metrics":["Accuracy (%)","MSE"],"first_row_in_archive_order":{"model":"GigaCheck (DN-DAB-DETR)","paper":"/paper/gigacheck-detecting-llm-generated-content","metrics":{"Accuracy (%)":"64.63","MSE":"1.51"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/gigacheck-detecting-llm-generated-content","title":"GigaCheck: Detecting LLM-generated Content","date":"2024-10-31","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/artificial-text-boundary-detection-with","title":"AI-generated text boundary detection with RoFT","date":"2023-11-14","rows_on_this_dataset":3,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":1,"samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":1,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}