Datasets › FEVER
FEVER (Fact Extraction and VERification)
FEVER is a publicly available dataset for fact extraction and verification against textual sources.
It consists of 185,445 claims manually verified against the introductory sections of Wikipedia pages and classified as SUPPORTED, REFUTED or NOTENOUGHINFO. For the first two classes, systems and annotators need to also return the combination of sentences forming the necessary evidence supporting or refuting the claim.
The claims were generated by human annotators extracting claims from Wikipedia and mutating them in a variety of ways, some of which were meaning-altering. The verification of each claim was conducted in a separate annotation process by annotators who were aware of the page but not the sentence from which original claim was extracted and thus in 31.75% of the claims more than one sentence was considered appropriate evidence. Claims require composition of evidence from multiple sentences in 16.82% of cases. Furthermore, in 12.15% of the claims, this evidence was taken from multiple pages.
Source: FEVER: a large-scale dataset for Fact Extraction and VERification
Benchmarks archive 2025-07-28
All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Question Answering | FEVER | CoA EM 68.9 | Chain-of-Action: Faithful and Multimodal Question... | MAGICS-LAB/Chain-of-Actions | 8 | Compare |
| Fact Verification | FEVER | ProoFVer-SB Accuracy 79.47 | ProoFVer: Natural Logic Theorem Proving for Fact Verification | krishnamrith12/proofver | 7 | Compare |
| Text Retrieval | FEVER | Lucene (BM25S) nDCG@10 63.8 | BM25S: Orders of magnitude faster lexical search via... | xhluca/bm25s +2 | 1 | Compare |
Papers archive 2025-07-28
11 shown of 11 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 498. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
Dataset loaders archive 2025-07-28
11 loaders as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- FEVER
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections