{"url":"/dataset/misinfo-general","name":"misinfo-general","full_name":null,"description_markdown":"We introduce `misinfo-general`, a benchmark dataset for evaluating misinformation models’ ability to perform out-of-distribution generalisation. Misinformation changes rapidly, much quicker than moderators can annotate at scale, resulting in a shift between the training and inference data distributions. As a result, misinformation models need to be able to perform out-of-distribution generalisation, an understudied problem in existing datasets.\r\n\r\nConstructed on top of the various NELA corpora ([2017](https://arxiv.org/abs/1803.10124), [2018](https://arxiv.org/abs/1904.01546), [2019](https://arxiv.org/abs/2003.08444), [2020](https://arxiv.org/abs/2102.04567), 2021, [2022](https://arxiv.org/abs/2203.05659)), `misinfo-general` is a large, diverse dataset consisting of news articles from reliable and unreliable publishers. Unlike NELA, we apply several rounds of deduplication and filtering to ensure all articles are of reasonable quality.\r\n\r\nWe use distant labelling to provide each publisher with rich metadata annotations. These annotations allow for simulating various generalisation splits that misinformation models are confronted with during deployment. We focus on 6 such splits-time, event, topic, publisher, political bias, misinformation type-but more are possible.\r\n\r\nBy releasing this dataset publicly, we hope to encourage future works that design misinformation models specifically with out-of-distribution generalisation in mind.","description_withheld":null,"homepage":"https://github.com/ioverho/misinfo-general","introduced_date":"2024-10-12","introduced_date_note":null,"introduced_by":{"paper":"/paper/yesterday-s-news-benchmarking-multi","title":"Yesterday's News: Benchmarking Multi-Dimensional Out-of-Distribution Generalisation of Misinformation Detection Models","first_author":"Ivo Verhoeven","url":null},"license":{"name":"CC BY-NC-ND 4.0","url":"https://creativecommons.org/licenses/by-nc-nd/4.0/"},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Misinformation","url":"/task/misinformation","datasets_with_task":"/datasets/task/misinformation"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["misinfo-general"],"data_loaders":[],"num_papers_in_archive":1,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}