Browse State-of-the-Art › Fact Verification
Fact Verification
129 papers with code · 3 benchmarks · 17 datasets archive 2025-07-28
Fact verification, also called "fact checking", is a process of verifying facts in natural text against a database of facts.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| KILT: FEVER (33 rows) | Re2G | Re2G: Retrieve, Rerank, Generate | code | Syntology ran 1 of 8 samples · 7 unverified | Compare |
| FEVER (7 rows) | ProoFVer-SB | ProoFVer: Natural Logic Theorem Proving for Fact Verification | code | — | Compare |
| DanFEVER (1 row) | DanFEVER XLM-RoBERTa Large | DanFEVER: claim verification dataset for Danish | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
17 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 129 papers with code (216 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
22 May 2020 18 repositories listed Syntology ran 4 of 6 samples · 2 unverifiedLarge pre-trained language models have been shown to store factual knowledge in their parameters, and achieve state-of-the-art results when fine-tuned on downstream NLP tasks.
-
6 Oct 2022 9 repositories listed Syntology ran 15 of 34 samples · 19 unverified · 5 pointer-only (licence)While large language models (LLMs) have demonstrated impressive capabilities across tasks in language understanding and interactive decision making, their abilities for reasoning (e.
-
17 Oct 2023 6 repositories listed Syntology ran 8 of 14 samples · 6 unverified · 3 pointer-only (licence)Our framework trains a single arbitrary LM that adaptively retrieves passages on-demand, and generates and reflects on retrieved passages and its own generations using special tokens, called reflection tokens.
-
20 Dec 2022 3 repositories listed Syntology ran 1 of 6 samples · 5 unverifiedGiven a query, HyDE first zero-shot instructs an instruction-following language model (e.
-
31 Dec 2020 3 repositories listedThis paper introduces the task of factual error correction: performing edits to a claim so that the generated rewrite is better supported by evidence.
-
4 Sep 2020 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)We test both task-specific and general baselines, evaluating downstream performance in addition to the ability of the models to provide provenance.
-
14 Aug 2019 3 repositories listedFact verification requires validating a claim in the context of evidence.
-
24 Feb 2025 2 repositories listedThis paper introduces Multi-Modal Retrieval-Augmented Generation (M^2RAG), a benchmark designed to evaluate the effectiveness of Multi-modal Large Language Models (MLLMs) in leveraging knowledge from multi-modal…
-
20 Feb 2024 2 repositories listedSimilar to the FEVER dataset, claims in the "Supports" and "Refutes" categories are also annotated with corresponding evidence sentences sourced from single or multiple pages in Chinese Wikipedia.
-
9 Jan 2024 2 repositories listed Syntology ran 6 of 8 samples · 2 unverifiedWe propose the Chain-of-Table framework, where tabular data is explicitly used in the reasoning chain as a proxy for intermediate thoughts.
-
14 Nov 2023 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)To alleviate these problems, we propose FILCO, a method that improves the quality of the context provided to the generator by (1) identifying useful context based on lexical and information-theoretic approaches, and (2)…
-
26 May 2023 2 repositories listedAlignScore is based on a general function of information alignment between two arbitrary text pieces.
-
20 Dec 2022 2 repositories listed Syntology ran 0 of 20 samples · 20 unverifiedHowever, constructing labeled data for complex reasoning tasks is labor intensive, and the quantity of annotated data remains insufficient to support the intricate demands of real-world applications.
-
16 Feb 2022 2 repositories listedMost of the existing debiasing methods often identify and weaken these samples with biased features (i.
-
3 Sep 2021 2 repositories listedWe introduce CREAK, a testbed for commonsense reasoning about entity knowledge, bridging fact-checking about entities (Harry Potter is a wizard and is skilled at riding a broomstick) with commonsense inferences (if…
-
5 Jul 2021 2 repositories listedClaims in FAVIQ are verified to be natural, contain little lexical bias, and require a complete understanding of the evidence for verification.
-
16 Dec 2020 2 repositories listedThis article investigates multilingual evidence retrieval and fact verification as a step to combat global disinformation, a first effort of this kind, to the best of our knowledge.
-
7 Oct 2020 2 repositories listed Syntology ran 0 of 5 samples · 5 unverifiedThe search can directly warn fake news posters and online users (e.
-
17 Sep 2019 2 repositories listed Syntology ran 6 of 17 samples · 11 unverifiedIn this work, we give general guidelines on system design for MRS by proposing a simple yet effective pipeline system with special consideration on hierarchical semantic retrieval at both paragraph and sentence level,…
-
13 Sep 2019 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)We experiment on large-scale natural language inference and fact verification benchmarks, evaluating on out-of-domain datasets that are specifically designed to assess the robustness of models against known biases in…
-
22 Jul 2019 2 repositories listedFact verification (FV) is a challenging task which requires to retrieve relevant evidence from plain text and use the evidence to verify given claims.
-
16 Nov 2018 2 repositories listed Syntology ran 4 of 5 samples · 1 unverifiedThe increasing concern with misinformation has stimulated research efforts on automatic fact checking.
-
8 Jul 2025 1 repository listedNumerical claims, statements involving quantities, comparisons, and temporal references, pose unique challenges for automated fact-checking systems.
-
16 Jun 2025 1 repository listedIn this study, we evaluate 12 pre-trained LLMs and one specialized fact-verifier, including frontier LLMs and open-weight reasoning LLMs, using a collection of examples from 14 fact-checking benchmarks.
-
10 Jun 2025 1 repository listedResults show that current models struggle with chart-based reasoning: even the best systems, such as Gemini 2.
-
2 Jun 2025 1 repository listedThe approach also achieves excellent performance on text-to-SQL tasks, reaching 68.
-
29 May 2025 1 repository listedWe develop and evaluate two post-training strategies to enable inference-time scaling: distillation from frontier model reasoning traces and reinforcement learning with verifiable rewards (RLVR).
-
2 Mar 2025 1 repository listedThe rise of misinformation, exacerbated by Large Language Models (LLMs) like GPT and Gemini, demands robust fact-checking solutions, especially for low-resource languages like Vietnamese.
-
24 Feb 2025 1 repository listed Syntology ran 0 of 3 samples · 3 unverifiedTabular data contains rich structural semantics and plays a crucial role in organizing and manipulating information.
-
20 Feb 2025 1 repository listedFact verification (FV) aims to assess the veracity of a claim based on relevant evidence.
Syntology lines on 13 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections