Browse State-of-the-Art › Multimodal Reasoning › Papers, page 3
Multimodal Reasoning
Papers archive 2025-07-28
archive papers tagged: 302 · with a code link: 138 · where Syntology ran a sample: 52 (43 with a run with no instrument failure, 9 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (52 of 302 tagged: 43 with a run with no instrument failure, 9 where every run was a failure of Syntology's instrument)
Page 3 of 4: papers 201 to 300 of 302, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Critique Before Thinking: Mitigating Hallucination through Rationale-Augmented Instruction Tuning12 May 2025 0 repositories listed
-
Skywork-VL Reward: An Effective Reward Model for Multimodal Understanding and Reasoning12 May 2025 0 repositories listed
-
Overview of the NLPCC 2025 Shared Task 4: Multi-modal, Multilingual, and Multi-hop Medical Instructional Video Question Answering Challenge11 May 2025 0 repositories listed
-
11 May 2025 0 repositories listed
-
Q-Heart: ECG Question Answering via Knowledge-Informed Multimodal LLMs7 May 2025 0 repositories listed
-
SToLa: Self-Adaptive Touch-Language Framework with Tactile Commonsense Reasoning in Open-Ended Scenarios7 May 2025 0 repositories listed
-
Advancing Conversational Diagnostic AI with Multimodal Reasoning6 May 2025 0 repositories listed
-
X-Reasoner: Towards Generalizable Reasoning Across Modalities and Domains6 May 2025 0 repositories listed
-
R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation4 May 2025 0 repositories listed
-
Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models30 Apr 2025 0 repositories listed
-
MultiMind: Enhancing Werewolf Agents with Multimodal Reasoning and Theory of Mind25 Apr 2025 0 repositories listed
-
GeoSense: Evaluating Identification and Application of Geometric Principles in Multimodal Reasoning17 Apr 2025 0 repositories listed
-
VLMGuard-R1: Proactive Safety Alignment for VLMs via Reasoning-Driven Prompt Optimization17 Apr 2025 0 repositories listed
-
Structured Graph Representations for Visual Narrative Reasoning: A Hierarchical Framework for Comics14 Apr 2025 0 repositories listed
-
SlowFastVAD: Video Anomaly Detection via Integrating Simple Detector and RAG-Enhanced Vision-Language Model14 Apr 2025 0 repositories listed
-
VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge14 Apr 2025 0 repositories listed
-
Draw with Thought: Unleashing Multimodal Reasoning for Scientific Diagram Generation13 Apr 2025 0 repositories listed
-
NoTeS-Bank: Benchmarking Neural Transcription and Search for Scientific Notes Understanding12 Apr 2025 0 repositories listed
-
VLMT: Vision-Language Multimodal Transformer for Multimodal Multi-hop Question Answering11 Apr 2025 0 repositories listed
-
MME-Unify: A Comprehensive Benchmark for Unified Multimodal Understanding and Generation Models4 Apr 2025 0 repositories listed
-
Why Reasoning Matters? A Survey of Advancements in Multimodal Reasoning (v1)4 Apr 2025 0 repositories listed
-
Agentic Multimodal AI for Hyperpersonalized B2B and B2C Advertising in Competitive Markets: An AI-Driven Competitive Advertising Framework1 Apr 2025 0 repositories listed
-
Evolutionary Prompt Optimization Discovers Emergent Multimodal Reasoning Strategies in Vision-Language Models30 Mar 2025 0 repositories listed
-
VisualQuest: A Diverse Image Dataset for Evaluating Visual Recognition in LLMs25 Mar 2025 0 repositories listed
-
Training-Free Personalization via Retrieval and Reasoning on Fingerprints24 Mar 2025 0 repositories listed
-
Mind with Eyes: from Language Reasoning to Multimodal Reasoning23 Mar 2025 0 repositories listed
-
Towards Agentic Recommender Systems in the Era of Multimodal Large Language Models20 Mar 2025 0 repositories listed
-
EfficientLLaVA:Generalizable Auto-Pruning for Large Vision-language Models19 Mar 2025 0 repositories listed
-
Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT Reasoning17 Mar 2025 0 repositories listed
-
MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification16 Mar 2025 0 repositories listed
-
VERIFY: A Benchmark of Visual Explanation and Reasoning for Investigating Multimodal Reasoning Fidelity14 Mar 2025 0 repositories listed
-
Chat-TS: Enhancing Multi-Modal Reasoning Over Time-Series and Natural Language Data13 Mar 2025 0 repositories listed
-
VisualPRM: An Effective Process Reward Model for Multimodal Reasoning13 Mar 2025 0 repositories listed
-
Seeing and Reasoning with Confidence: Supercharging Multimodal LLMs with an Uncertainty-Aware Agentic Framework11 Mar 2025 0 repositories listed
-
Integrating Chain-of-Thought for Multimodal Alignment: A Study on 3D Vision-Language Learning8 Mar 2025 0 repositories listed
-
COSINT-Agent: A Knowledge-Driven Multimodal Agent for Chinese Open Source Intelligence5 Mar 2025 0 repositories listed
-
All-in-one: Understanding and Generation in Multimodal Reasoning with the MAIA Benchmark24 Feb 2025 0 repositories listed
-
Shakti-VLMs: Scalable Vision-Language Models for Enterprise AI24 Feb 2025 0 repositories listed
-
Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models22 Feb 2025 0 repositories listed
-
Exploring Advanced Techniques for Visual Question Answering: A Comprehensive Comparison20 Feb 2025 0 repositories listed
-
CutPaste&Find: Efficient Multimodal Hallucination Detector with Visual-aid Knowledge Base18 Feb 2025 0 repositories listed
-
EnigmaEval: A Benchmark of Long Multimodal Reasoning Challenges13 Feb 2025 0 repositories listed
-
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency13 Feb 2025 0 repositories listed
-
A Generative Framework for Bidirectional Image-Report Understanding in Chest Radiography9 Feb 2025 0 repositories listed
-
Boosting Multimodal Reasoning with MCTS-Automated Structured Thinking4 Feb 2025 0 repositories listed
-
Mitigating Object Hallucinations in Large Vision-Language Models via Attention Calibration4 Feb 2025 0 repositories listed
-
Position: Empowering Time Series Reasoning with Multimodal LLMs3 Feb 2025 0 repositories listed
-
Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark9 Jan 2025 0 repositories listed
-
DRIVINGVQA: Analyzing Visual Chain-of-Thought Reasoning of Vision Language Models in Real-World Scenarios with Driving Theory Tests8 Jan 2025 0 repositories listed
-
EfficientLLaVA: Generalizable Auto-Pruning for Large Vision-language Models1 Jan 2025 0 repositories listed
-
Diving into Self-Evolving Training for Multimodal Reasoning23 Dec 2024 0 repositories listed
-
FiVL: A Framework for Improved Vision-Language Alignment19 Dec 2024 0 repositories listed
-
Progressive Multimodal Reasoning via Active Retrieval19 Dec 2024 0 repositories listed
-
Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence18 Dec 2024 0 repositories listed
-
A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges16 Dec 2024 0 repositories listed
-
Leveraging Retrieval-Augmented Tags for Large Vision-Language Understanding in Complex Scenes16 Dec 2024 0 repositories listed
-
Optimizing Vision-Language Interactions Through Decoder-Only Models14 Dec 2024 0 repositories listed
-
EVLM: Self-Reflective Multimodal Reasoning for Cross-Dimensional Visual Editing13 Dec 2024 0 repositories listed
-
Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning27 Nov 2024 0 repositories listed
-
Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous Driving20 Nov 2024 0 repositories listed
-
Motion-Grounded Video Reasoning: Understanding and Perceiving Motion at Pixel Level15 Nov 2024 0 repositories listed
-
Learning to Ground VLMs without Forgetting14 Oct 2024 0 repositories listed
-
An X-Ray Is Worth 15 Features: Sparse Autoencoders for Interpretable Radiology Report Generation4 Oct 2024 0 repositories listed
-
NL-Eye: Abductive NLI for Images3 Oct 2024 0 repositories listed
-
Deep Learning and Machine Learning, Advancing Big Data Analytics and Management: Unveiling AI's Potential Through Tools, Techniques, and Applications2 Oct 2024 0 repositories listed
-
Proof of Thought : Neurosymbolic Program Synthesis allows Robust and Interpretable Reasoning25 Sep 2024 0 repositories listed
-
NVLM: Open Frontier-Class Multimodal LLMs17 Sep 2024 0 repositories listed
-
Knowledge-Aware Reasoning over Multimodal Semi-structured Tables25 Aug 2024 0 repositories listed
-
Towards Holistic Disease Risk Prediction using Small Language Models13 Aug 2024 0 repositories listed
-
User-in-the-loop Evaluation of Multimodal LLMs for Activity Assistance4 Aug 2024 0 repositories listed
-
On scalable oversight with weak LLMs judging strong LLMs5 Jul 2024 0 repositories listed
-
Improving Multi-Agent Debate with Sparse Communication Topology17 Jun 2024 0 repositories listed
-
POEM: Interactive Prompt Optimization for Enhancing Multimodal Reasoning of Large Language Models6 Jun 2024 0 repositories listed
-
Multimodal Reasoning with Multimodal Knowledge Graph4 Jun 2024 0 repositories listed
-
Retrieval Meets Reasoning: Even High-school Textbook Knowledge Benefits Multimodal Reasoning31 May 2024 0 repositories listed
-
22 May 2024 0 repositories listed
-
Inquire, Interact, and Integrate: A Proactive Agent Collaborative Framework for Zero-Shot Multimodal Medical Reasoning19 May 2024 0 repositories listed
-
AccidentBlip: Agent of Accident Warning based on MA-former18 Apr 2024 0 repositories listed
-
Closed-Loop Open-Vocabulary Mobile Manipulation with GPT-4V16 Apr 2024 0 repositories listed
-
Text Is MASS: Modeling as Stochastic Embedding for Text-Video Retrieval26 Mar 2024 0 repositories listed
-
25 Feb 2024 0 repositories listed
-
Exploring Failure Cases in Multimodal Reasoning About Physical Dynamics24 Feb 2024 0 repositories listed
-
BBA: Bi-Modal Behavioral Alignment for Reasoning with Large Vision-Language Models21 Feb 2024 0 repositories listed
-
Question Aware Vision Transformer for Multimodal Reasoning8 Feb 2024 0 repositories listed
-
Hidden flaws behind expert-level accuracy of multimodal GPT-4 vision in medicine16 Jan 2024 0 repositories listed
-
Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning10 Jan 2024 0 repositories listed
-
Assessing GPT4-V on Structured Reasoning Tasks13 Dec 2023 0 repositories listed
-
DDCoT: Duty-Distinct Chain-of-Thought Prompting for Multimodal Reasoning in Language Models25 Oct 2023 0 repositories listed
-
Personality-aware Human-centric Multimodal Reasoning: A New Task, Dataset and Baselines5 Apr 2023 0 repositories listed
-
AutoFraudNet: A Multimodal Network to Detect Fraud in the Auto Insurance Industry15 Jan 2023 0 repositories listed
-
Deep Neural Networks for Visual Reasoning24 Sep 2022 0 repositories listed
-
Reducing the Vision and Language Bias for Temporal Sentence Grounding27 Jul 2022 0 repositories listed
-
DisinfoMeme: A Multimodal Dataset for Detecting Meme Intentionally Spreading Out Disinformation25 May 2022 0 repositories listed
-
Learning from Inside: Self-driven Siamese Sampling and Reasoning for Video Question Answering1 Dec 2021 0 repositories listed
-
Improving Pre-trained Vision-and-Language Embeddings for Phrase Grounding1 Nov 2021 0 repositories listed
-
TxT: Crossmodal End-to-End Learning with Transformers9 Sep 2021 0 repositories listed
-
C³: Compositional Counterfactual Contrastive Learning for Video-grounded Dialogues16 Jun 2021 0 repositories listed
-
Premise-based Multimodal Reasoning: Conditional Inference on Joint Textual and Visual Clues15 May 2021 0 repositories listed
-
DOC2PPT: Automatic Presentation Slides Generation from Scientific Documents28 Jan 2021 0 repositories listed
-
Sound2Sight: Generating Visual Dynamics from Sound and Context23 Jul 2020 0 repositories listed