Home › Datasets › task › Misinformation

Misinformation datasets

archive 2025-07-28

44 datasets carry the task tag "Misinformation" (the task itself: Misinformation), ordered by the archive's paper count. Page 1 of 1: 44 shown of 44. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Misinformation datasets 1–44 of 44

The Hateful Memes data set is a multimodal dataset for hateful meme detection (image + text) that contains 10,000+ new multimodal examples created by Facebook AI.
177 papers · 3 benchmarks
LIAR is a publicly available dataset for fake news detection.
130 papers · 1 benchmark
DFDC (Deepfake Detection Challenge)
The DFDC (Deepfake Detection Challenge) is a dataset for deepface detection consisting of more than 100,000 videos.
126 papers · 1 benchmark
A new challenge set for multimodal classification, focusing on detecting hate speech in multimodal memes.
48 papers · 0 benchmarks
CoAID include diverse COVID-19 healthcare misinformation, including fake news on websites and social platforms, along with users' social engagement about such news.
41 papers · 0 benchmarks
A new publicly available dataset for verification of climate change-related claims.
38 papers · 1 benchmark
FakeNewsNet is collected from two fact-checking websites: GossipCop and PolitiFact containing news contents with labels annotated by professional journalists and experts, along with social context information.
30 papers · 0 benchmarks
UPFD (User Preference-aware Fake News Detection)
For benchmarking, please refer to its variant UPFD-POL and UPFD-GOS.
13 papers · 0 benchmarks
A large-scale curated dataset of over 152 million tweets, growing daily, related to COVID-19 chatter generated from January 1st to April 4th at the time of writing.
10 papers · 0 benchmarks
For LIAR-RAW, we extended the public dataset LIAR-PLUS (Alhindi et al., 2018) with relevant raw reports, containing fine-grained claims from Politifact.
9 papers · 0 benchmarks
MEIR (Multimodal Entity Image Repurposing)
MEIR is a substantially challenging dataset over that which has been previously available to support research into image repurposing detection.
9 papers · 0 benchmarks
NELA-GT-2018 is a dataset for the study of misinformation that consists of 713k articles collected between 02/2018-11/2018.
9 papers · 0 benchmarks
ArCOV-19 is an Arabic COVID-19 Twitter dataset that covers the period from 27th of January till 30th of April 2020.
6 papers · 0 benchmarks
For RAWFC, we constructed it from scratch by collecting the claims from Snopes and relevant raw reports by retrieving claim keywords.
6 papers · 1 benchmark
Chinese dataset on COVID-19 misinformation.
5 papers · 0 benchmarks
MuMiN is a misinformation graph dataset containing rich social media data (tweets, replies, users, images, articles, hashtags), spanning 21 million tweets belonging to 26 thousand Twitter threads, each of which have been semantically…
5 papers · 0 benchmarks
NELA-GT-2019 is an updated version of the NELA-GT-2018 dataset.
5 papers · 0 benchmarks
ArCOV19-Rumors is an Arabic COVID-19 Twitter dataset for misinformation detection composed of tweets containing claims from 27th January till the end of April 2020.
4 papers · 0 benchmarks
CoVERT (A Corpus of Fact-checked Biomedical COVID-19 Tweets)
CoVERT is a fact-checked corpus of tweets with a focus on the domain of biomedicine and COVID-19-related (mis)information.
4 papers · 0 benchmarks
GeoCoV19 is a large-scale Twitter dataset containing more than 524 million multilingual tweets.
4 papers · 0 benchmarks
NELA-GT-2020 is an updated version of the NELA-GT-2019 dataset.
4 papers · 0 benchmarks
Some Like it Hoax is a fake news detection dataset consisting of 15,500 Facebook posts and 909,236 users.
4 papers · 0 benchmarks
COVID-Q consists of COVID-19 questions which have been annotated into a broad category (e.g.
3 papers · 0 benchmarks
CoVaxLies v1 includes 17 known Misinformation Targets (MisTs) found on Twitter about the covid-19 vaccines.
3 papers · 0 benchmarks
Stanceosaurus is a corpus of 28,033 tweets in English, Hindi, and Arabic annotated with stance towards 251 misinformation claims.
3 papers · 0 benchmarks
Detecting out-of-context media, such as "mis-captioned" images on Twitter, is a relevant problem, especially in domains of high public significance.
3 papers · 0 benchmarks
UPFD-GOS (User Preference-aware Fake News Detection)
The Gossipcop variant of the UPFD dataset for benchmarking.
3 papers · 1 benchmark
CoVaxFrames includes 113 Vaccine Hesitancy Framings found on Twitter about the COVID-19 vaccines.
2 papers · 0 benchmarks
CoVaxLies v2 includes 47 Misinformation Targets (MisTs) found on Twitter about the COVID-19 vaccines.
2 papers · 0 benchmarks
Covid-HeRA is a dataset for health risk assessment and severity-informed decision making in the presence of COVID19 misinformation.
2 papers · 0 benchmarks
MMVax-Stance includes 113 Vaccine Hesitancy Framings found on Twitter about the COVID-19 vaccines.
2 papers · 0 benchmarks
ThreatGram 101 - Extreme Telegram Data (ThreatGram 101 - Extreme Telegram Replies Data with Threat Levels)
Data 1: Raw and Unlabeled; 2 million unlabeled replies from 17 Telegram channels.
2 papers · 1 benchmark
UPFD-POL (User Preference-aware Fake News Detection)
The PolitiFact variant of the UPFD dataset for benchmarking.
2 papers · 1 benchmark
This paper analyses two hitherto unstudied sites sharing state-backed disinformation, Reliable Recent News (rrn.world) and WarOnFakes (waronfakes.com), which publish content in Arabic, Chinese, English, French, German, and Spanish.
1 paper · 0 benchmarks
HAVOC (Harmful Abstractions and Violations in Open Completions Benchmark)
measure the toxicity generated by language models across input severity and harm categories, by creating a new benchmark of open ended prefixes.
1 paper · 0 benchmarks
HpVaxFrames includes 64 Vaccine Hesitancy Framings found on Twitter about the HPV vaccines.
1 paper · 0 benchmarks
MIPD (Manipulation and Intention In a Novel Corpus of Polish Disinformation)
A novel corpus of 15,356 Polish web articles, including articles identified as containing disinformation.
1 paper · 0 benchmarks
This dataset of medical misinformation was collected and is published by Kempelen Institute of Intelligent Technologies (KInIT).
1 paper · 0 benchmarks
NAIST COVID is a multilingual dataset of social media posts related to COVID-19, consisting of microblogs in English and Japanese from Twitter and those in Chinese from Weibo.
1 paper · 0 benchmarks
NELA-GT-2021 is the fourth installment of the NELA-GT datasets, NELA-GT-2021.
1 paper · 0 benchmarks
Combines CoVaxFrames and HpVaxFrames into a unified dataset of 113 Vaccine Hesitancy Framings found on Twitter about the COVID-19 vaccines and 64 Vaccine Hesitancy Framings found on Twitter about the HPV vaccines.
1 paper · 0 benchmarks
A Natural Language Resource for Learning to Recognize Misinformation about the COVID-19 and HPV Vaccines.
1 paper · 0 benchmarks
We introduce misinfo-general, a benchmark dataset for evaluating misinformation models’ ability to perform out-of-distribution generalisation.
1 paper · 0 benchmarks
CIDII Dataset (Correct Information and Disinformation about Islamic Issues)
The CIDII dataset is a binary classification, consisting of two classes of correct information and disinformation related to Islamic issues.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.