Datasets › Poly-FEVER

Poly-FEVER

Introduced by Hanzhi Zhang et al. in Poly-FEVER: A Multilingual Fact Verification Benchmark for Hallucination Detection in Large Language Models19 Mar 2025 archive 2025-07-28

Poly-FEVER is a multilingual fact verification benchmark designed to evaluate hallucination detection in large language models (LLMs). It extends three widely used fact-checking datasets—FEVER, Climate-FEVER, and SciFact—by translating claims into 11 languages, enabling cross-linguistic analysis of LLM performance.

Poly-FEVER consists of 77,973 factual claims with binary labels (SUPPORTS or REFUTES), making it suitable for benchmarking multilingual hallucination detection. The dataset covers various domains, including Arts, Science, Politics, and History.

Funded by [optional]: Google Cloud Translation Language(s) (NLP): English(en), Mandarin Chinese (zh-CN), Hindi (hi), Arabic (ar), Bengali (bn), Japanese (ja), Korean (ko), Tamil (ta), Thai (th), Georgian (ka), and Amharic (am)

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset; the archive counts 1 paper for it but never published that list.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

No task tagged in the archive.

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Poly-FEVER

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections