Browse State-of-the-Art › Multimodal Reasoning
Multimodal Reasoning
138 papers with code · 3 benchmarks · 11 datasets archive 2025-07-28
Reasoning over multimodal inputs.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| REBUS (8 rows) | GPT-4V | REBUS: A Robust Evaluation Benchmark of Understanding Symbols | code | — | Compare |
| MATH-V (4 rows) | GPT4V | Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset | code | — | Compare |
| AlgoPuzzleVQA (1 row) | GPT-4 | Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
11 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 138 papers with code (302 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
24 May 2025 3 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedWe showcase the utility of this dataset through: 1) improving infographic chart understanding via fine-tuning, 2) benchmarking code generation for infographic charts, and 3) enabling example-based infographic chart…
-
1 Sep 2021 3 repositories listed Syntology ran 4 of 13 samples · 9 unverified · 12 pointer-only (licence)Scaling Visual Question Answering (VQA) to the open-domain and multi-hop nature of web searches, requires fundamental advances in visual representation learning, knowledge aggregation, and language generation.
-
7 Apr 2020 3 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 2 pointer-only (licence)The recently proposed SNLI-VE corpus for recognising visual-textual entailment is a large, real-world dataset for fine-grained multimodal reasoning.
-
8 Jul 2025 2 repositories listed Syntology ran 3 of 12 samples · 9 unverified · 12 pointer-only (licence)The strong performance of Skywork-R1V3 primarily stems from our elaborate post-training RL framework, which effectively activates and enhances the model's reasoning ability, without the need for additional continue…
-
28 May 2025 2 repositories listedRecent advancements in large language models (LLMs) have demonstrated impressive chain-of-thought reasoning capabilities, with reinforcement learning (RL) playing a crucial role in this progress.
-
20 May 2025 2 repositories listed Syntology ran 9 of 21 samples · 12 unverifiedUnifying multimodal understanding and generation has shown impressive capabilities in cutting-edge proprietary systems.
-
19 May 2025 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedWhile Multimodal Large Language Models (MLLMs) have achieved impressive progress in vision-language understanding, they still struggle with complex multi-step reasoning, often producing logically inconsistent or…
-
19 May 2025 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)While their performances on speech and audio-processing tasks are extensively studied, their reasoning abilities remain underexplored.
-
10 Apr 2025 2 repositories listed Syntology ran 2 of 12 samples · 10 unverifiedBy combining these two techniques, our model, VL-Rethinker, advances state-of-the-art scores on MathVista, MathVerse to achieve 80.
-
3 Feb 2025 2 repositories listedOur results reveal that o-[n] series, particularly later iterations like o3 and o4-mini, significantly outperform the GPT-[n] series and show strong scalability in multimodal reasoning.
-
15 Nov 2024 2 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedLarge language models have demonstrated substantial advancements in reasoning capabilities, particularly through inference-time scaling, as illustrated by models such as OpenAI's o1.
-
24 Oct 2024 2 repositories listed Syntology ran 0 of 13 samples · 13 unverified · 13 pointer-only (licence)Specifically, we employ text-based synthesizing techniques to construct chart-plotting code and produce ReachQA, a dataset containing 3k reasoning-intensive charts and 20k Q&A pairs to enhance both recognition and…
-
2 Aug 2024 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedMotivated by this, we introduce MuChoMusic, a benchmark for evaluating music understanding in multimodal language models focused on audio.
-
20 Mar 2024 2 repositories listed Syntology ran 13 of 14 samples · 1 unverifiedTo diagnose the reasoning challenges in large multimodal models, we progressively guide the models with our ground truth reasoning explanations for visual perception, inductive reasoning, and deductive reasoning.
-
6 Mar 2024 2 repositories listedWe present a new dataset, AlgoPuzzleVQA designed to challenge and evaluate the capabilities of multimodal language models in solving algorithmic puzzles that necessitate both visual understanding, language…
-
13 Oct 2023 2 repositories listedConsequently, our work complements research on the performance of MLLMs in multimodal comprehension tasks, achieving a more comprehensive and holistic evaluation of MLLMs.
-
26 May 2023 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Therefore, we propose Graph-of-Thought (GoT) reasoning, which models human thought processes not only as a chain but also as a graph.
-
1 Oct 2022 2 repositories listed Syntology ran 7 of 27 samples · 20 unverifiedAnalogical reasoning is fundamental to human cognition and holds an important place in various fields.
-
2 Nov 2016 2 repositories listedWe propose Dual Attention Networks (DANs) which jointly leverage visual and textual attention mechanisms to capture fine-grained interplay between vision and language.
-
6 Jul 2025 1 repository listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Recent advances in vision-language-action (VLA) models have shown promise in integrating image generation with action prediction to improve generalization and reasoning in robot manipulation.
-
1 Jul 2025 1 repository listed Syntology ran 0 of 11 samples · 11 unverifiedIn a comprehensive evaluation across 28 public benchmarks, our model outperforms Qwen2.
-
30 Jun 2025 1 repository listed(1) We establish the foundational principles of the think with image paradigm and its three-stage framework.
-
26 Jun 2025 1 repository listedWith the rapid evolution of multimodal large language models, the capacity to deeply understand and interpret human intentions has emerged as a critical capability, which demands detailed and thoughtful reasoning.
-
20 Jun 2025 1 repository listed Syntology ran 0 of 17 samples · 17 unverified · 17 pointer-only (licence)Vision-language models (VLMs) excel at multimodal understanding, yet their text-only decoding forces them to verbalize visual reasoning, limiting performance on tasks that demand visual imagination.
-
16 Jun 2025 1 repository listedRecent advancements in large language models (LLMs) have witnessed a surge in the development of advanced reasoning paradigms, which are now being integrated into multimodal large language models (MLLMs).
-
11 Jun 2025 1 repository listedTo address the limitations, we propose drawing to reason in space, a novel paradigm that enables LVLMs to reason through elementary drawing operations in the visual space.
-
9 Jun 2025 1 repository listed Syntology ran 0 of 12 samples · 12 unverifiedDeveloping generalizable reasoning capabilities in multimodal large language models (MLLMs) remains challenging.
-
9 Jun 2025 1 repository listed(2) The open-source WeThink dataset containing over 120K multimodal QA pairs with annotated reasoning paths, curated from 18 diverse dataset sources and covering various question domains.
-
8 Jun 2025 1 repository listedSpecifically, we propose a Spatial Token Fusion (STF) method to learn compact vision tokens for short vision token sequence, where spatial-adjacent tokens are fused into one.
-
5 Jun 2025 1 repository listedDespite rapid advances in vision-language models (VLMs), current benchmarks for multimodal reasoning fall short in three key dimensions.
Syntology lines on 18 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections