Browse State-of-the-Art › Logical Reasoning
Logical Reasoning
330 papers with code · 10 benchmarks · 18 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
10 leaderboard tables shown for this task, 10 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
18 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
22 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 330 papers with code (747 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
22 Oct 2020 9 repositories listed Syntology ran 8 of 14 samples · 6 unverified · 1 pointer-only (licence)Logical operations are performed in the embedding space by neural operators over the probabilistic embeddings.
-
5 Apr 2022 7 repositories listed Syntology ran 30 of 37 samples · 7 unverifiedTo further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model PaLM.
-
27 Nov 2023 5 repositories listed Syntology ran 6 of 14 samples · 8 unverified · 1 pointer-only (licence)We introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning.
-
28 Jun 2024 4 repositories listed Syntology ran 1 of 4 samples · 3 unverified · 1 pointer-only (licence)We propose a novel persona-driven data synthesis methodology that leverages various perspectives within a large language model (LLM) to create diverse synthetic data.
-
24 May 2022 4 repositories listed Syntology ran 0 of 4 samples · 4 unverified · 1 pointer-only (licence)Pretrained large language models (LLMs) are widely used in many sub-fields of natural language processing (NLP) and generally known as excellent few-shot learners with task-specific exemplars.
-
4 Nov 2024 3 repositories listedIn this paper, we introduce Hunyuan-Large, which is currently the largest open-source Transformer-based mixture of experts model, with a total of 389 billion parameters and 52 billion activation parameters, capable of…
-
19 Jan 2024 3 repositories listed Syntology ran 1 of 4 samples · 3 unverified · 4 pointer-only (licence)In this paper, we introduce the notion of knowledge fusion for LLMs, aimed at combining the capabilities of existing LLMs and transferring them into a single LLM.
-
11 Aug 2023 3 repositories listed Syntology ran 0 of 2 samples · 2 unverifiedWe rethink this and adopt a well-grounded set of deduction rules based on formal logic theory, which can derive any other deduction rules when combined in a multistep way.
-
22 Oct 2022 3 repositories listedTo this end, we propose a comprehensive logical reasoning explanation form.
-
8 Dec 2021 3 repositories listedLanguage modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.
-
20 Aug 2020 3 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedBoth reasoning and generalization ability are important for prediction tasks such as recommender systems, where reasoning provides strong connection between user history and target items for accurate prediction, and…
-
16 May 2020 3 repositories listedExisting Collaborative Filtering (CF) methods are mostly designed based on the idea of matching, i.
-
29 May 2019 3 repositories listed Syntology ran 4 of 5 samples · 1 unverifiedWe demonstrate that by integrating this solver into end-to-end learning systems, we can learn the logical structure of challenging problems in a minimally supervised fashion.
-
16 Mar 2018 3 repositories listed Syntology ran 5 of 5 samples · 0 unverifiedCOG is much simpler than the general problem of video analysis, yet it addresses many of the problems relating to visual and logical reasoning and memory -- problems that remain challenging for modern deep learning…
-
19 Apr 2025 2 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedRecent works have begun exploring reasoning in GUI tasks with encouraging results.
-
25 Feb 2025 2 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Reasoning is a fundamental capability of large language models (LLMs), enabling them to comprehend, analyze, and solve complex problems.
-
19 Nov 2024 2 repositories listedLarge language models (LLMs) are capable of solving a wide range of tasks, yet they have struggled with reasoning.
-
15 Nov 2024 2 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedLarge language models have demonstrated substantial advancements in reasoning capabilities, particularly through inference-time scaling, as illustrated by models such as OpenAI's o1.
-
7 Oct 2024 2 repositories listedWhile the performance of LLMs on GSM8K has significantly improved in recent years, it remains unclear whether their mathematical reasoning capabilities have genuinely advanced, raising questions about the reliability of…
-
16 Jul 2024 2 repositories listedTo address these limitations, we introduce NeedleBench, a synthetic framework for assessing retrieval and reasoning performance in bilingual long-context tasks with adaptive context lengths.
-
19 Feb 2024 2 repositories listedWe further characterize the game-theoretic properties of LLMs, such as equilibrium and Pareto Efficiency in repeated games.
-
12 Dec 2023 2 repositories listed Syntology ran 9 of 17 samples · 8 unverifiedSGLang consists of a frontend language and a runtime.
-
26 Oct 2023 2 repositories listedMaking neural visual generative models controllable by logical reasoning systems is promising for improving faithfulness, transparency, and generalizability.
-
23 May 2023 2 repositories listed Syntology ran 2 of 8 samples · 6 unverifiedExisting efforts to improve logical reasoning ability of language models have predominantly relied on supervised fine-tuning, hindering generalization to new domains and/or tasks.
-
3 Apr 2023 2 repositories listedWe propose the polytuplet loss function, which forces prioritization of learning the relative correctness of answer choices over learning the true accuracy of each choice.
-
30 Mar 2023 2 repositories listedThe use of NLP in the realm of financial technology is broad and complex, with applications ranging from sentiment analysis and named entity recognition to question answering.
-
28 Mar 2023 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedIn this paper, we propose LEAP, a novel system that uses language models to perform multi-step logical reasoning and incorporates explicit planning into the inference procedure.
-
19 Dec 2022 2 repositories listedReasoning, as an essential ability for complex problem-solving, can provide back-end support for various real-world applications, such as medical diagnosis, negotiation, etc.
-
6 Dec 2022 2 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 4 pointer-only (licence)Naturally, we also present a unified multi-task Geometric Transformer framework, Geoformer, to tackle calculation and proving problems simultaneously in the form of sequence generation, which finally shows the reasoning…
-
16 Oct 2022 2 repositories listedTo verify this hypothesis, we manually construct a set of counterfactual samples, which modify the original logical forms to generate counterfactual logical forms with rarely co-occurred table headers and logical…
Syntology lines on 17 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections