Home › Datasets › task › Fake News Detection
Fake News Detection datasets
archive 2025-07-28
30 datasets carry the task tag "Fake News Detection" (the task itself: Fake News Detection), ordered by the archive's paper count. Page 1 of 1: 30 shown of 30. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Fake News Detection datasets 1–30 of 30
LIAR is a publicly available dataset for fake news detection.
130 papers · 1 benchmark
RealNews is a large corpus of news articles from Common Crawl.
80 papers · 0 benchmarks
The Weibo NER dataset is a Chinese Named Entity Recognition dataset drawn from the social media website Sina Weibo.
52 papers · 2 benchmarks
FakeNewsNet is collected from two fact-checking websites: GossipCop and PolitiFact containing news contents with labels annotated by professional journalists and experts, along with social context information.
30 papers · 0 benchmarks
Fact-checking (FC) articles which contains pairs (multimodal tweet and a FC-article) from snopes.com.
22 papers · 1 benchmark
Fact-checking (FC) articles which contains pairs (multimodal tweet and a FC-article) from politifact.com.
20 papers · 1 benchmark
FNC-1 (Fake News Challenge Stage 1)
FNC-1 was designed as a stance detection dataset and it contains 75,385 labeled headline and article pairs.
19 papers · 2 benchmarks
Weibo21 is a benchmark of fake news dataset for multi-domain fake news detection (MFND) with domain label annotated, which consists of 4,488 fake news and 4,640 real news from 9 different domains.
18 papers · 0 benchmarks
Fakeddit is a novel multimodal dataset for fake news detection consisting of over 1 million samples from multiple categories of fake news.
15 papers · 0 benchmarks
UPFD (User Preference-aware Fake News Detection)
For benchmarking, please refer to its variant UPFD-POL and UPFD-GOS.
13 papers · 0 benchmarks
Along with COVID-19 pandemic we are also fighting an infodemic'.
12 papers · 1 benchmark
MM-COVID (Multilingual and Multidimensional COVID-19 Fake News Data Repository)
MM-COVID is a dataset for fake news detection related to COVID-19.
12 papers · 0 benchmarks
For LIAR-RAW, we extended the public dataset LIAR-PLUS (Alhindi et al., 2018) with relevant raw reports, containing fine-grained claims from Politifact.
9 papers · 0 benchmarks
NELA-GT-2018 is a dataset for the study of misinformation that consists of 713k articles collected between 02/2018-11/2018.
9 papers · 0 benchmarks
An annotated dataset of ~50K news that can be used for building automated fake news detection systems for a low resource language like Bangla.
7 papers · 0 benchmarks
For RAWFC, we constructed it from scratch by collecting the claims from Snopes and relevant raw reports by retrieving claim keywords.
6 papers · 1 benchmark
MuMiN is a misinformation graph dataset containing rich social media data (tweets, replies, users, images, articles, hashtags), spanning 21 million tweets belonging to 26 thousand Twitter threads, each of which have been semantically…
5 papers · 0 benchmarks
NELA-GT-2019 is an updated version of the NELA-GT-2018 dataset.
5 papers · 0 benchmarks
The LIAR dataset has been widely followed by fake news detection researchers since its release, and along with a great deal of research, the community has provided a variety of feedback on the dataset to improve it.
4 papers · 1 benchmark
NELA-GT-2020 is an updated version of the NELA-GT-2019 dataset.
4 papers · 0 benchmarks
Some Like it Hoax is a fake news detection dataset consisting of 15,500 Facebook posts and 909,236 users.
4 papers · 0 benchmarks
UPFD-GOS (User Preference-aware Fake News Detection)
The Gossipcop variant of the UPFD dataset for benchmarking.
3 papers · 1 benchmark
AraCOVID19-MFH (AraCOVID19-MFH: Arabic COVID-19 Multi-label Fake News and Hate Speech Detection Dataset)
AraCOVID19-MFH is a manually annotated multi-label Arabic COVID-19 fake news and hate speech detection dataset.
2 papers · 0 benchmarks
A Dataset to Identify Manipulated Social Media News in Bangla We construct a publicly available Bangla dataset of 800 news-related social media items that are annotated as manipulated or not relative to 500 reference news articles.
2 papers · 0 benchmarks
News SEO Dataset (Detection and Discovery of Misinformation Sources using Attributed Webgraphs)
Search Engine Optimization (SEO) attributes provide strong signals for predicting news site reliability.
2 papers · 0 benchmarks
UPFD-POL (User Preference-aware Fake News Detection)
The PolitiFact variant of the UPFD dataset for benchmarking.
2 papers · 1 benchmark
Expertly-curated benchmark dataset for fake news detection in Filipino.
1 paper · 0 benchmarks
MIPD (Manipulation and Intention In a Novel Corpus of Polish Disinformation)
A novel corpus of 15,356 Polish web articles, including articles identified as containing disinformation.
1 paper · 0 benchmarks
The task addresses the problem of the appearance and propagation of posts that share misleading multimedia content (images or video).
1 paper · 0 benchmarks
CIDII Dataset (Correct Information and Disinformation about Islamic Issues)
The CIDII dataset is a binary classification, consisting of two classes of correct information and disinformation related to Islamic issues.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.