Datasets › FEVER

FEVER (Fact Extraction and VERification)

Introduced by James Thorne et al. in FEVER: a large-scale dataset for Fact Extraction and VERification14 Mar 2018 archive 2025-07-28

FEVER is a publicly available dataset for fact extraction and verification against textual sources.

It consists of 185,445 claims manually verified against the introductory sections of Wikipedia pages and classified as SUPPORTED, REFUTED or NOTENOUGHINFO. For the first two classes, systems and annotators need to also return the combination of sentences forming the necessary evidence supporting or refuting the claim.

The claims were generated by human annotators extracting claims from Wikipedia and mutating them in a variety of ways, some of which were meaning-altering. The verification of each claim was conducted in a separate annotation process by annotators who were aware of the page but not the sentence from which original claim was extracted and thus in 31.75% of the claims more than one sentence was considered appropriate evidence. Claims require composition of evidence from multiple sentences in 16.82% of cases. Furthermore, in 12.15% of the claims, this evidence was taken from multiple pages.

Source: FEVER: a large-scale dataset for Fact Extraction and VERification

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

11 shown of 11 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 498. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
BM25S: Orders of magnitude faster lexical search via eager sparse scoring 3 1 4 Jul 2024 ran 10 of 22 samples (12 unverified)
Chain-of-Action: Faithful and Multimodal Question Answering through Large Language Models 1 5 26 Mar 2024 ran 8 of 10 samples (2 unverified)
DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines 3 1 5 Oct 2023 ran 3 of 7 samples (4 unverified)
Measuring and Narrowing the Compositionality Gap in Language Models 1 1 7 Oct 2022 not harvested
Paragraph-based Transformer Pre-training for Multi-Sentence Inference 1 2 2 May 2022 not harvested
ProoFVer: Natural Logic Theorem Proving for Fact Verification 1 1 25 Aug 2021 not harvested
Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks 18 1 22 May 2020 ran 4 of 6 samples (2 unverified)
Fine-grained Fact Verification with Kernel Graph Attention Network 1 1 22 Oct 2019 ran 4 of 5 samples (1 unverified; 3 pointer-only for licence)
Reasoning Over Semantic-Level Graph for Fact Checking 0 1 9 Sep 2019 not harvested
GEAR: Graph-based Evidence Aggregating and Reasoning for Fact Verification 2 1 22 Jul 2019 not harvested
Language Models are Unsupervised Multitask Learners 21 1 14 Feb 2019 not harvested

Dataset loaders archive 2025-07-28

11 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • FEVER

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections