Browse State-of-the-Art › Arithmetic Reasoning › Papers, page 2
Arithmetic Reasoning
Papers archive 2025-07-28
archive papers tagged: 175 · with a code link: 112 · where Syntology ran a sample: 64 (53 with a run with no instrument failure, 11 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (64 of 175 tagged: 53 with a run with no instrument failure, 11 where every run was a failure of Syntology's instrument)
Page 2 of 2: papers 101 to 175 of 175, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
15 Feb 2023 1 repository listed
-
8 Feb 2023 1 repository listed
-
31 Jan 2023 1 repository listed
-
19 Dec 2022 1 repository listed Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
3 Nov 2022 1 repository listed
-
28 Oct 2022 1 repository listed
-
12 Oct 2022 1 repository listed
-
29 Jun 2022 1 repository listed
-
28 May 2022 1 repository listed Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples)
-
21 May 2022 1 repository listed
-
25 Oct 2021 1 repository listed Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
10 May 2021 1 repository listed Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
FinLMM-R1: Enhancing Financial Reasoning in LMM through Scalable Data and Reward Design16 Jun 2025 0 repositories listed
-
Learning-at-Criticality in Large Language Models for Quantum Field Theory and Beyond4 Jun 2025 0 repositories listed
-
VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL29 May 2025 0 repositories listed
-
Joint Flashback Adaptation for Forgetting-Resistant Instruction Tuning21 May 2025 0 repositories listed
-
Tokenization Constraints in LLMs: A Study of Symbolic and Arithmetic Reasoning Limits20 May 2025 0 repositories listed
-
Fact-Consistency Evaluation of Text-to-SQL Generation for Business Intelligence Using Exaone 3.530 Apr 2025 0 repositories listed
-
ThoughtProbe: Classifier-Guided Thought Space Exploration Leveraging LLM Intrinsic Reasoning9 Apr 2025 0 repositories listed
-
Your Language Model May Think Too Rigidly: Achieving Reasoning Consistency with Symmetry-Enhanced Training25 Feb 2025 0 repositories listed
-
The Lottery LLM Hypothesis, Rethinking What Abilities Should LLM Compression Preserve?24 Feb 2025 0 repositories listed
-
Inference-Time Computations for LLM Reasoning and Planning: A Benchmark and Insights18 Feb 2025 0 repositories listed
-
On Representational Dissociation of Language and Arithmetic in Large Language Models17 Feb 2025 0 repositories listed
-
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding17 Feb 2025 0 repositories listed
-
Can LLMs Maintain Fundamental Abilities under KV Cache Compression?4 Feb 2025 0 repositories listed
-
CLoQ: Enhancing Fine-Tuning of Quantized LLMs via Calibrated LoRA Initialization30 Jan 2025 0 repositories listed
-
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training28 Jan 2025 0 repositories listed
-
DoTA: Weight-Decomposed Tensor Adaptation for Large Language Models30 Dec 2024 0 repositories listed
-
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning23 Dec 2024 0 repositories listed
-
Why We Build Local Large Language Models: An Observational Analysis from 35 Japanese and Multilingual LLMs19 Dec 2024 0 repositories listed
-
Hint Marginalization for Improved Reasoning in Large Language Models17 Dec 2024 0 repositories listed
-
GaLore+: Boosting Low-Rank Adaptation for LLMs with Cross-Head Projection15 Dec 2024 0 repositories listed
-
S²FT: Efficient, Scalable and Generalizable LLM Fine-tuning by Structured Sparsity9 Dec 2024 0 repositories listed
-
Think-to-Talk or Talk-to-Think? When LLMs Come Up with an Answer in Multi-Step Arithmetic Reasoning2 Dec 2024 0 repositories listed
-
PERFT: Parameter-Efficient Routed Fine-Tuning for Mixture-of-Expert Model12 Nov 2024 0 repositories listed
-
Think Beyond Size: Adaptive Prompting for More Effective Reasoning10 Oct 2024 0 repositories listed
-
Unlocking Structured Thinking in Language Models with Cognitive Prompting3 Oct 2024 0 repositories listed
-
Small Language Models are Equation Reasoners19 Sep 2024 0 repositories listed
-
Relating the Seemingly Unrelated: Principled Understanding of Generalization for Generative Models in Arithmetic Reasoning Tasks25 Jul 2024 0 repositories listed
-
Leveraging LLM Reasoning Enhances Personalized Recommender Systems22 Jul 2024 0 repositories listed
-
Fine-Tuning and Prompt Optimization: Two Great Steps that Work Better Together15 Jul 2024 0 repositories listed
-
Arithmetic Reasoning with LLM: Prolog Generation & Permutation28 May 2024 0 repositories listed
-
Large Language Models Can Self-Correct with Key Condition Verification23 May 2024 0 repositories listed
-
Skin-in-the-Game: Decision Making via Multi-Stakeholder Alignment in LLMs21 May 2024 0 repositories listed
-
Enabling High-Sparsity Foundational Llama Models with Efficient Pretraining and Deployment6 May 2024 0 repositories listed
-
4 Mar 2024 0 repositories listed
-
SymBa: Symbolic Backward Chaining for Structured Natural Language Reasoning20 Feb 2024 0 repositories listed
-
Evaluating LLMs' Mathematical Reasoning in Financial Document Question Answering17 Feb 2024 0 repositories listed
-
16 Feb 2024 0 repositories listed
-
Exploring Group and Symmetry Principles in Large Language Models9 Feb 2024 0 repositories listed
-
9 Feb 2024 0 repositories listed
-
Evaluating Gender Bias in Large Language Models via Chain-of-Thought Prompting28 Jan 2024 0 repositories listed
-
Large Language Models are Null-Shot Learners16 Jan 2024 0 repositories listed
-
14 Dec 2023 0 repositories listed
-
14 Dec 2023 0 repositories listed
-
18 Nov 2023 0 repositories listed
-
14 Nov 2023 0 repositories listed
-
Prompt Sketching for Large Language Models8 Nov 2023 0 repositories listed
-
11 Oct 2023 0 repositories listed
-
11 Jul 2023 0 repositories listed
-
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes23 Jun 2023 0 repositories listed
-
DiversiGATE: A Comprehensive Framework for Reliable Large Language Models22 Jun 2023 0 repositories listed
-
Code Prompting: a Neural Symbolic Method for Complex Reasoning in Large Language Models29 May 2023 0 repositories listed
-
RCOT: Detecting and Rectifying Factual Inconsistency in Reasoning by Reversing Chain-of-Thought19 May 2023 0 repositories listed
-
Hint of Thought prompting: an explainable and zero-shot approach to reasoning tasks with LLMs19 May 2023 0 repositories listed
-
Self-Evaluation Guided Beam Search for Reasoning1 May 2023 0 repositories listed
-
When do you need Chain-of-Thought Prompting for ChatGPT?6 Apr 2023 0 repositories listed
-
25 Nov 2022 0 repositories listed
-
20 Oct 2022 0 repositories listed
-
20 Oct 2022 0 repositories listed
-
20 Oct 2022 0 repositories listed
-
Neural-Symbolic Recursive Machine for Systematic Generalization4 Oct 2022 0 repositories listed
-
19 Jun 2022 0 repositories listed
-
6 Jun 2022 0 repositories listed
-
12 Apr 2022 0 repositories listed
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.