Home › Datasets › task › Fact Checking
Fact Checking datasets
archive 2025-07-28
12 datasets carry the task tag "Fact Checking" (the task itself: Fact Checking), ordered by the archive's paper count. Page 1 of 1: 12 shown of 12. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Fact Checking datasets 1–12 of 12
BEIR (Benchmarking IR) is a heterogeneous benchmark containing different information retrieval (IR) tasks.
311 papers · 10 benchmarks
VitaminC (Fact Verification with Contrastive Evidence)
The VitaminC dataset contains more than 450,000 claim-evidence pairs for fact verification and factual consistent generation.
42 papers · 0 benchmarks
AVeriTeC (AVeriTeC: A Dataset for Real-world Claim Verification with Evidence from the Web)
AVeriTeC (Automated Verification of Textual Claims) is a dataset of 4568 real-world claims covering fact-checks by 50 different organizations.
17 papers · 1 benchmark
A large-scale dataset that consists of 21,184 claims, where each claim is assigned a truthfulness label and ruling statement, with 58,523 pieces of evidence in the form of text and images.
10 papers · 0 benchmarks
ataset format Each row in the dataset splits represents one instance and contains the following tab-separated columns: articleid - article id corresponding to the id of the claim in the LIAR dataset statement - the text of the claim author…
7 papers · 0 benchmarks
CoVERT (A Corpus of Fact-checked Biomedical COVID-19 Tweets)
CoVERT is a fact-checked corpus of tweets with a focus on the domain of biomedicine and COVID-19-related (mis)information.
4 papers · 0 benchmarks
The LIAR dataset has been widely followed by fake news detection researchers since its release, and along with a great deal of research, the community has provided a variety of feedback on the dataset to improve it.
4 papers · 1 benchmark
Stanceosaurus is a corpus of 28,033 tweets in English, Hindi, and Arabic annotated with stance towards 251 misinformation claims.
3 papers · 0 benchmarks
- CFEVER is a Chinese Fact Extraction and VERification dataset published at AAAI 2024.
1 paper · 0 benchmarks
Spiced is a paraphrase dataset of scientific findings annotated for degree of information change.
1 paper · 0 benchmarks
PANACEA (PANACEA dataset - Heterogeneous COVID-19 Claims)
The peer-reviewed publication for this dataset has been presented in the 2022 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), and can be accessed here:…
0 papers · 0 benchmarks
STVD-FC is the largest public dataset on the political content analysis and fact-checking tasks.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.