{"url":"/dataset/halueval","name":"HaluEval","full_name":null,"description_markdown":"**HaluEval** is a **large-scale hallucination evaluation benchmark** designed for **Large Language Models (LLMs)**. It provides a comprehensive collection of **generated and human-annotated hallucinated samples** to evaluate the performance of LLMs in recognizing hallucinations¹².\r\n\r\nHere are the key details about the **HaluEval dataset**:\r\n\r\n1. **Purpose and Overview**:\r\n   - **Purpose**: HaluEval aims to understand the types of content and the extent to which LLMs are prone to hallucinate.\r\n   - **Content**: It includes both **general user queries** with ChatGPT responses and **task-specific examples** from three tasks: question answering, knowledge-grounded dialogue, and text summarization.\r\n   - **Data Sources**:\r\n     - For general user queries, HaluEval adopts the **52K instruction tuning dataset** from Alpaca.\r\n     - Task-specific examples are generated based on existing task datasets (e.g., HotpotQA, OpenDialKG, CNN/Daily Mail) as seed data.\r\n\r\n2. **Data Composition**:\r\n   - **General User Queries**:\r\n     - 5,000 user queries paired with ChatGPT responses.\r\n     - Queries are selected based on low-similarity responses to identify potential hallucinations.\r\n   - **Task-Specific Examples**:\r\n     - 30,000 examples from three tasks:\r\n       - **Question Answering**: Based on HotpotQA as seed data.\r\n       - **Knowledge-Grounded Dialogue**: Based on OpenDialKG as seed data.\r\n       - **Text Summarization**: Based on CNN/Daily Mail as seed data.\r\n\r\n3. **Data Release**:\r\n   - The dataset contains **35,000 generated and human-annotated hallucinated samples** used in experiments.\r\n   - JSON files include:\r\n     - `qa_data.json`: Hallucinated QA samples.\r\n     - `dialogue_data.json`: Hallucinated dialogue samples.\r\n     - `summarization_data.json`: Hallucinated summarization samples.\r\n     - `general_data.json`: Human-annotated ChatGPT responses to general user queries.\r\n\r\nSource: Conversation with Bing, 3/17/2024\r\n(1) HaluEval: A Hallucination Evaluation Benchmark for LLMs. https://github.com/RUCAIBox/HaluEval.\r\n(2) jzjiao/halueval-sft · Datasets at Hugging Face. https://huggingface.co/datasets/jzjiao/halueval-sft.\r\n(3) HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large .... https://aclanthology.org/2023.emnlp-main.397/.\r\n(4) HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large .... https://arxiv.org/abs/2305.11747.\r\n(5) undefined. https://github.com/RUCAIBox/HaluEval%29.","description_withheld":null,"homepage":"https://github.com/RUCAIBox/HaluEval","introduced_date":"2023-05-19","introduced_date_note":null,"introduced_by":{"paper":"/paper/helma-a-large-scale-hallucination-evaluation","title":"HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models","first_author":"Junyi Li","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["HaluEval"],"data_loaders":[],"num_papers_in_archive":68,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}