Browse State-of-the-Art › StrategyQA
StrategyQA
20 papers with code · 0 benchmarks · 0 datasets archive 2025-07-28
StrategyQA aims to measure the ability of models to answer questions that require multi-step implicit reasoning.
Source: BIG-bench
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
No dataset record in the archive lists this task.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
20 shown of 20 papers with code (40 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
5 Apr 2022 7 repositories listed Syntology ran 30 of 37 samples · 7 unverifiedTo further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model PaLM.
-
12 Aug 2024 3 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedThis paper introduces rStar, a self-play mutual reasoning approach that significantly improves reasoning capabilities of small language models (SLMs) without fine-tuning or superior models.
-
30 Sep 2023 3 repositories listedTo improve planning with LLMs, we propose an agentic architecture, the Modular Agentic Planner (MAP), in which planning is accomplished via the recurrent interaction of the specialized modules mentioned above, each…
-
21 Mar 2022 3 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedChain-of-thought prompting combined with pre-trained large language models has achieved encouraging results on complex reasoning tasks.
-
8 Dec 2021 3 repositories listedLanguage modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.
-
29 Mar 2022 2 repositories listed Syntology ran 8 of 11 samples · 3 unverified · 4 pointer-only (licence)We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget.
-
4 Mar 2025 1 repository listedLarge Language Models (LLMs) are increasingly being used in real-world applications.
-
26 Feb 2025 1 repository listedOur results show that voting protocols improve performance by 13.
-
9 Dec 2024 1 repository listedChain of Thought (CoT) was introduced in recent research as a method for improving step-by-step reasoning in Large Language Models.
-
7 Oct 2024 1 repository listed Syntology ran 6 of 7 samples · 1 unverifiedFurthermore, we demonstrate that training a verifier on valid rationales significantly improves its ability to distinguish valid and flawed rationale.
-
23 May 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)In MoE, each token in the input sequence activates a different subset of experts determined by a routing mechanism.
-
3 Mar 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)In this work, we seek a novel KGQA dataset that supports commonsense reasoning and focuses on long-tail entities (e.
-
21 Feb 2024 1 repository listed Syntology ran 6 of 10 samples · 4 unverified · 10 pointer-only (licence)We propose a straightforward approach called Distillation Contrastive Decoding (DCD) to enhance the reasoning capabilities of Large Language Models (LLMs) during inference.
-
19 Jan 2024 1 repository listed Syntology ran 7 of 8 samples · 1 unverified · 8 pointer-only (licence)Self-consistency (SC) has been a widely used decoding strategy for chain-of-thought reasoning.
-
6 Nov 2023 1 repository listed Syntology ran 6 of 10 samples · 4 unverified · 10 pointer-only (licence)Results on five difficult question-answering datasets StrategyQA, QuaRel, OpenBookQA, NumerSense and QASC show that not only does MaRio improve task accuracy, but it also improves the self-rationalization quality of…
-
2 Aug 2023 1 repository listedWe equip a smaller Language Model to generalise to answering challenging compositional questions that have not been seen in training.
-
28 May 2023 1 repository listed Syntology ran 2 of 6 samples · 4 unverifiedLarge Language Models (LLMs) have shown promising performance in knowledge-intensive reasoning tasks that require a compound understanding of knowledge.
-
19 Dec 2022 1 repository listedThis paper proposes a question-answering system that can answer questions whose supporting evidence is spread over multiple (potentially long) documents.
-
1 Dec 2022 1 repository listedIn this work, we propose an alternative reasoning scheme, Socratic CoT, that learns a decomposition of the original problem into a sequence of subproblems and uses it to guide the intermediate reasoning steps.
-
6 Jan 2021 1 repository listedA key limitation in current datasets for multi-hop reasoning is that the required steps for answering the question are mentioned in it explicitly.
Syntology lines on 11 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections