Browse State-of-the-Art › Fact Checking
Fact Checking
297 papers with code · 7 benchmarks · 12 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
7 leaderboard tables shown for this task, 7 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| SciFact (BEIR) (5 rows) | monoT5-3B | No Parameter Left Behind: How Distillation and Model Size Affect... | code | — | Compare |
| CLIMATE-FEVER (BEIR) (4 rows) | SGPT-BE-5.8B | SGPT: GPT Sentence Embeddings for Semantic Search | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
| FEVER (BEIR) (4 rows) | monoT5-3B | No Parameter Left Behind: How Distillation and Model Size Affect... | code | — | Compare |
| AVeriTeC (3 rows) | HerO | HerO at AVeriTeC: The Herd of Open Large Language Models for... | code | — | Compare |
| ^(#!@#)(()))****** (1 row) | Abc | MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents | code | Syntology ran 2 of 8 samples · 6 unverified | Compare |
| CDCD (1 row) | MA-CIN | Self-Supervised Claim Identification for Automated Fact Checking | code | — | Compare |
| LIAR2 (1 row) | FDHN | An Enhanced Fake News Detection System With Fuzzy Deep Learning | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
12 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
5 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 297 papers with code (669 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
1 May 2017 11 repositories listedIn this paper, we present liar: a new, publicly available dataset for fake news detection.
-
11 Jul 2017 9 repositories listedIdentifying public misinformation is a complicated and challenging task.
-
16 Dec 2021 6 repositories listed Syntology ran 1 of 5 samples · 4 unverifiedIn this work, we explore the limits of contrastive learning as a way to train unsupervised dense retrievers and show that it leads to strong performance in various retrieval settings.
-
19 May 2021 6 repositories listedThe proliferation of fake news, i.
-
9 May 2024 4 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)To mitigate these issues, we propose OpenFactCheck, a unified framework for building customized automatic fact-checking systems, benchmarking their accuracy, evaluating factuality of LLMs, and verifying claims in a…
-
28 Oct 2019 4 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedCurrently used metrics for assessing summarization algorithms do not account for whether summaries are factually consistent with source documents.
-
10 Feb 2019 4 repositories listed Syntology ran 2 of 12 samples · 10 unverified · 1 pointer-only (licence)One of the main reasons is that often the interpretation of the news requires the knowledge of political or social context or 'common sense', which current NLP algorithms are still missing.
-
25 Jul 2023 3 repositories listed Syntology ran 4 of 14 samples · 10 unverifiedWith the above challenges in mind, in this paper, we propose FacTool, a task and domain agnostic framework for detecting factual errors of texts generated by large language models (e.
-
22 May 2023 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Existing datasets for automated fact-checking have substantial limitations, such as relying on artificial claims, lacking annotations for evidence and intermediate reasoning, or including evidence published after the…
-
8 Dec 2021 3 repositories listedLanguage modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.
-
2 Dec 2021 3 repositories listed Syntology ran 1 of 13 samples · 12 unverifiedOur approach outperforms two competitive baselines on three scientific claim verification datasets, with particularly strong performance in zero / few-shot domain adaptation experiments.
-
17 Apr 2021 3 repositories listed Syntology ran 3 of 4 samples · 1 unverifiedTo address this, and to facilitate researchers to broadly evaluate the effectiveness of their models, we introduce Benchmarking-IR (BEIR), a robust and heterogeneous evaluation benchmark for information retrieval.
-
16 Apr 2021 3 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedWe present KnowledgeEditor, a method which can be used to edit this knowledge and, thus, fix 'bugs' or unexpected predictions without the need for expensive re-training or fine-tuning.
-
31 Dec 2020 3 repositories listedThis paper introduces the task of factual error correction: performing edits to a claim so that the generated rewrite is better supported by evidence.
-
7 Sep 2020 3 repositories listedWhile misinformation and disinformation have been thriving in social media for years, with the emergence of the COVID-19 pandemic, the political and the health misinformation merged, thus elevating the problem to a…
-
4 Sep 2020 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)We test both task-specific and general baselines, evaluating downstream performance in addition to the ability of the models to provide provenance.
-
26 Mar 2020 3 repositories listedThe analysis is presented and updated on a publically accessible dashboard (https://usc-melady.
-
21 Jan 2020 3 repositories listedFinally, the lab offers a fifth task that asks to predict the check-worthiness of the claims made in English political debates and speeches.
-
30 Sep 2019 3 repositories listedThis is a challenging constrained generation task, as the output must be consistent with the new information and fit into the rest of the existing document.
-
8 Mar 2018 3 repositories listedCommunity Question Answering (cQA) forums are very popular nowadays, as they represent effective means for communities around particular topics to share information.
-
3 Feb 2025 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedDebunking them requires (1) providing the true context of the image and (2) checking the veracity of the image's caption.
-
6 Aug 2024 2 repositories listedThe increased use of large language models (LLMs) across a variety of real-world applications calls for automatic tools to check the factual accuracy of their outputs, as LLMs often hallucinate.
-
16 Jul 2024 2 repositories listedOur model specifies the expected behavior of each operator with a high-quality gold algorithm, and we develop an optimization framework that reduces cost, while providing accuracy guarantees with respect to a gold…
-
5 Jun 2024 2 repositories listedUnlike previous fallacy detection datasets, Missci (i) focuses on implicit fallacies between the relevant content of the cited publication and the inaccurate claim, and (ii) requires models to verbalize the fallacious…
-
16 Apr 2024 2 repositories listed Syntology ran 2 of 8 samples · 6 unverifiedWe release LLM-AggreFact, code for data synthesis, and models.
-
18 Mar 2024 2 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 5 pointer-only (licence)Streaming text generation has become a common way of increasing the responsiveness of language model powered applications, such as chat assistants.
-
20 Feb 2024 2 repositories listedSimilar to the FEVER dataset, claims in the "Supports" and "Refutes" categories are also annotated with corresponding evidence sentences sourced from single or multiple pages in Chinese Wikipedia.
-
15 Nov 2023 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedThe increased use of large language models (LLMs) across a variety of real-world applications calls for mechanisms to verify the factual accuracy of their outputs.
-
4 Jun 2023 2 repositories listedWe run the first systematic evaluation of pre-trained language models for Bulgarian, comparing and contrasting results across the nine tasks in the benchmark.
-
22 May 2023 2 repositories listed Syntology ran 2 of 6 samples · 4 unverifiedFact-checking real-world claims often requires collecting multiple pieces of evidence and applying complex multi-step reasoning.
Syntology lines on 15 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections