Datasets › WikiBio GPT-3 Hallucination Dataset
WikiBio GPT-3 Hallucination Dataset
The WikiBio GPT-3 Hallucination Dataset is a benchmark dataset used for hallucination detection. It is based on Wikipedia biographies (WikiBio) and is specifically designed to evaluate the factuality of text generated by large language models like GPT-3¹². Here are some key details about this dataset:
- Dataset Source: Wikipedia biographies (WikiBio)
- Task: Text classification
- Language: English
- Size Categories: Less than 1,000 samples
- License: Creative Commons Attribution-ShareAlike 3.0 (cc-by-sa-3.0)
(1) potsawee/wiki_bio_gpt3_hallucination · Datasets at Hugging Face. https://huggingface.co/datasets/potsawee/wiki_bio_gpt3_hallucination. (2) [2303.08896] SelfCheckGPT: Zero-Resource Black-Box Hallucination .... https://arxiv.org/abs/2303.08896. (3) AIトラストと、対話型生成AIにおける富士通のAIトラスト技術 : 富士通. https://www.fujitsu.com/jp/about/research/article/202312-ai-trust-technologies.html. (4) README.md · potsawee/wiki_bio_gpt3_hallucination at main - Hugging Face. https://huggingface.co/datasets/potsawee/wiki_bio_gpt3_hallucination/blob/main/README.md. (5) undefined. https://github.com/potsawee/selfcheckgpt.
Benchmarks archive 2025-07-28
No leaderboard in the archive resolves to this dataset.
Papers archive 2025-07-28
No paper in the archive has a leaderboard row on this dataset; the archive counts 1 paper for it but never published that list.
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
No task tagged in the archive.
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
No modality tagged.
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- WikiBio GPT-3 Hallucination Dataset
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections