Browse State-of-the-Art › Visual Reasoning › Papers, page 5
Visual Reasoning
Papers archive 2025-07-28
archive papers tagged: 698 · with a code link: 356 · where Syntology ran a sample: 165 (130 with a run with no instrument failure, 35 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (165 of 698 tagged: 130 with a run with no instrument failure, 35 where every run was a failure of Syntology's instrument)
Page 5 of 7: papers 401 to 500 of 698, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models19 May 2025 0 repositories listed
-
Human-Aligned Bench: Fine-Grained Assessment of Reasoning Ability in MLLMs vs. Humans16 May 2025 0 repositories listed
-
VLM Q-Learning: Aligning Vision-Language Models for Interactive Decision-Making6 May 2025 0 repositories listed
-
A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law5 May 2025 0 repositories listed
-
Localizing Before Answering: A Hallucination Evaluation Benchmark for Grounded Medical Multimodal LLMs30 Apr 2025 0 repositories listed
-
NORA: A Small Open-Sourced Generalist Vision Language Action Model for Embodied Tasks28 Apr 2025 0 repositories listed
-
A Comprehensive Survey of Knowledge-Based Vision Question Answering Systems: The Lifecycle of Knowledge in Visual Reasoning Task24 Apr 2025 0 repositories listed
-
LongPerceptualThoughts: Distilling System-2 Reasoning for System-1 Perception21 Apr 2025 0 repositories listed
-
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models21 Apr 2025 0 repositories listed
-
Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?18 Apr 2025 0 repositories listed
-
LVLM_CSP: Accelerating Large Vision Language Models via Clustering, Scattering, and Pruning for Reasoning Segmentation15 Apr 2025 0 repositories listed
-
Visual Language Models show widespread visual deficits on neuropsychological tests15 Apr 2025 0 repositories listed
-
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography14 Apr 2025 0 repositories listed
-
VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge14 Apr 2025 0 repositories listed
-
On Data Synthesis and Post-training for Visual Abstract Reasoning2 Apr 2025 0 repositories listed
-
GenVP: Generating Visual Puzzles with Contrastive Hierarchical VAEs30 Mar 2025 0 repositories listed
-
DWIM: Towards Tool-aware Visual Reasoning via Discrepancy-aware Workflow Generation & Instruct-Masking Tuning25 Mar 2025 0 repositories listed
-
RGB-Th-Bench: A Dense benchmark for Visual-Thermal Understanding of Vision Language Models25 Mar 2025 0 repositories listed
-
Neuro-Symbolic Scene Graph Conditioning for Synthetic Image Dataset Generation21 Mar 2025 0 repositories listed
-
Chain of Functions: A Programmatic Pipeline for Fine-Grained Chart Reasoning Data20 Mar 2025 0 repositories listed
-
From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibration17 Mar 2025 0 repositories listed
-
VERIFY: A Benchmark of Visual Explanation and Reasoning for Investigating Multimodal Reasoning Fidelity14 Mar 2025 0 repositories listed
-
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems13 Mar 2025 0 repositories listed
-
Does Acceleration Cause Hidden Instability in Vision Language Models? Uncovering Instance-Level Divergence Through a Large-Scale Empirical Study9 Mar 2025 0 repositories listed
-
Poisoned-MRAG: Knowledge Poisoning Attacks to Multimodal Retrieval Augmented Generation8 Mar 2025 0 repositories listed
-
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression6 Mar 2025 0 repositories listed
-
Towards Visual Discrimination and Reasoning of Real-World Physical Dynamics: Physics-Grounded Anomaly Detection5 Mar 2025 0 repositories listed
-
EXCLAIM: An Explainable Cross-Modal Agentic System for Misinformation Detection with Hierarchical Retrieval1 Mar 2025 0 repositories listed
-
M-LLM Based Video Frame Selection for Efficient Video Understanding27 Feb 2025 0 repositories listed
-
MMSciBench: Benchmarking Language Models on Multimodal Scientific Problems27 Feb 2025 0 repositories listed
-
End-to-End Chart Summarization via Visual Chain-of-Thought in Vision-Language Models24 Feb 2025 0 repositories listed
-
Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI24 Feb 2025 0 repositories listed
-
VisFactor: Benchmarking Fundamental Visual Cognition in Multimodal Large Language Models23 Feb 2025 0 repositories listed
-
Visual Reasoning Evaluation of Grok, Deepseek Janus, Gemini, Qwen, Mistral, and ChatGPT23 Feb 2025 0 repositories listed
-
Chitrarth: Bridging Vision and Language for a Billion People21 Feb 2025 0 repositories listed
-
Do we Really Need Visual Instructions? Towards Visual Instruction-Free Fine-tuning for Large Vision-Language Models17 Feb 2025 0 repositories listed
-
Learning to Stop Overthinking at Test Time16 Feb 2025 0 repositories listed
-
MET-Bench: Multimodal Entity Tracking for Evaluating the Limitations of Vision-Language and Reasoning Models15 Feb 2025 0 repositories listed
-
ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models13 Feb 2025 0 repositories listed
-
Visual Agentic AI for Spatial Reasoning with a Dynamic API10 Feb 2025 0 repositories listed
-
Boosting Multimodal Reasoning with MCTS-Automated Structured Thinking4 Feb 2025 0 repositories listed
-
Integrating LMM Planners and 3D Skill Policies for Generalizable Manipulation30 Jan 2025 0 repositories listed
-
Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models30 Jan 2025 0 repositories listed
-
3D-MoE: A Mixture-of-Experts Multi-modal LLM for 3D Vision and Pose Diffusion via Rectified Flow28 Jan 2025 0 repositories listed
-
23 Jan 2025 0 repositories listed
-
Systematic Abductive Reasoning via Diverse Relation Representations in Vector-symbolic Architecture21 Jan 2025 0 repositories listed
-
MAPS: Advancing Multi-Modal Reasoning in Expert-Level Physical Science18 Jan 2025 0 repositories listed
-
CityLoc: 6DoF Pose Distributional Localization for Text Descriptions in Large-Scale Scenes with Gaussian Representation15 Jan 2025 0 repositories listed
-
DRIVINGVQA: Analyzing Visual Chain-of-Thought Reasoning of Vision Language Models in Real-World Scenarios with Driving Theory Tests8 Jan 2025 0 repositories listed
-
From Code to Compliance: Assessing ChatGPT's Utility in Designing an Accessible Webpage -- A Case Study7 Jan 2025 0 repositories listed
-
LogicAD: Explainable Anomaly Detection via VLM-based Text Feature Extraction3 Jan 2025 0 repositories listed
-
Language-Guided Salient Object Ranking1 Jan 2025 0 repositories listed
-
Probing Visual Language Priors in VLMs31 Dec 2024 0 repositories listed
-
Slow Perception: Let's Perceive Geometric Figures Step-by-step30 Dec 2024 0 repositories listed
-
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities21 Dec 2024 0 repositories listed
-
EarthDial: Turning Multi-sensory Earth Observations to Interactive Dialogues19 Dec 2024 0 repositories listed
-
ViUniT: Visual Unit Tests for More Robust Visual Programming12 Dec 2024 0 repositories listed
-
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models4 Dec 2024 0 repositories listed
-
VISCO: Benchmarking Fine-Grained Critique and Correction Towards Self-Improvement in Visual Reasoning3 Dec 2024 0 repositories listed
-
Abductive Symbolic Solver on Abstraction and Reasoning Corpus27 Nov 2024 0 repositories listed
-
Enhancing Visual Reasoning with Autonomous Imagination in Multimodal Large Language Models27 Nov 2024 0 repositories listed
-
Beyond Visual Understanding: Introducing PARROT-360V for Vision Language Model Benchmarking20 Nov 2024 0 repositories listed
-
Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios20 Nov 2024 0 repositories listed
-
Automated 3D Physical Simulation of Open-world Scene with Gaussian Splatting19 Nov 2024 0 repositories listed
-
Bootstrapping Top-down Information for Self-modulating Slot Attention4 Nov 2024 0 repositories listed
-
Reasoning Limitations of Multimodal Large Language Models. A case study of Bongard Problems2 Nov 2024 0 repositories listed
-
Replace-then-Perturb: Targeted Adversarial Attacks With Visual Reasoning for Vision-Language Models1 Nov 2024 0 repositories listed
-
VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning30 Oct 2024 0 repositories listed
-
Improving Generalization in Visual Reasoning via Self-Ensemble28 Oct 2024 0 repositories listed
-
ProReason: Multi-Modal Proactive Reasoning with Decoupled Eyesight and Wisdom18 Oct 2024 0 repositories listed
-
ForgeryGPT: Multimodal Large Language Model For Explainable Image Forgery Detection and Localization14 Oct 2024 0 repositories listed
-
10 Oct 2024 0 repositories listed
-
Transformers Utilization in Chart Understanding: A Review of Recent Advances & Future Trends5 Oct 2024 0 repositories listed
-
Advancing Object Detection in Transportation with Multimodal Large Language Models (MLLMs): A Comprehensive Review and Empirical Testing26 Sep 2024 0 repositories listed
-
GSON: A Group-based Social Navigation Framework with Large Multimodal Model26 Sep 2024 0 repositories listed
-
Enhancing Advanced Visual Reasoning Ability of Large Language Models21 Sep 2024 0 repositories listed
-
Impact of ML Optimization Tactics on Greener Pre-Trained ML Models19 Sep 2024 0 repositories listed
-
What Makes a Maze Look Like a Maze?12 Sep 2024 0 repositories listed
-
Critical Features Tracking on Triangulated Irregular Networks by a Scale-Space Method10 Sep 2024 0 repositories listed
-
MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct9 Sep 2024 0 repositories listed
-
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis27 Aug 2024 0 repositories listed
-
Compromising Embodied Agents with Contextual Backdoor Attacks6 Aug 2024 0 repositories listed
-
ExoViP: Step-by-step Verification and Exploration with Exoskeleton Modules for Compositional Visual Reasoning5 Aug 2024 0 repositories listed
-
Chat2Layout: Interactive 3D Furniture Layout with a Multimodal LLM31 Jul 2024 0 repositories listed
-
Pyramid Coder: Hierarchical Code Generator for Compositional Visual Question Answering30 Jul 2024 0 repositories listed
-
Take A Step Back: Rethinking the Two Stages in Visual Reasoning29 Jul 2024 0 repositories listed
-
Can VLMs be used on videos for action recognition? LLMs are Visual Reasoning Coordinators20 Jul 2024 0 repositories listed
-
I Know About "Up"! Enhancing Spatial Reasoning in Visual Language Models Through 3D Reconstruction19 Jul 2024 0 repositories listed
-
Open-World Visual Reasoning by a Neuro-Symbolic Program of Zero-Shot Symbols18 Jul 2024 0 repositories listed
-
X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs18 Jul 2024 0 repositories listed
-
SwitchCIT: Switching for Continual Instruction Tuning16 Jul 2024 0 repositories listed
-
NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models15 Jul 2024 0 repositories listed
-
Affordance-Guided Reinforcement Learning via Visual Prompting14 Jul 2024 0 repositories listed
-
NODE-Adapter: Neural Ordinary Differential Equations for Better Vision-Language Reasoning11 Jul 2024 0 repositories listed
-
MMRo: Are Multimodal LLMs Eligible as the Brain for In-Home Robotics?28 Jun 2024 0 repositories listed
-
Disentangling Knowledge-based and Visual Reasoning by Question Decomposition in KB-VQA27 Jun 2024 0 repositories listed
-
Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration24 Jun 2024 0 repositories listed
-
Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities20 Jun 2024 0 repositories listed
-
GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs19 Jun 2024 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.