Browse State-of-the-Art › Visual Question Answering › Papers, page 11
Visual Question Answering
Papers archive 2025-07-28
archive papers tagged: 2,177 · with a code link: 1,042 · where Syntology ran a sample: 378 (308 with a run with no instrument failure, 70 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (378 of 2,177 tagged: 308 with a run with no instrument failure, 70 where every run was a failure of Syntology's instrument)
Page 11 of 22: papers 1,001 to 1,100 of 2,177, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
10 Jul 2018 1 repository listed Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
9 Jul 2018 1 repository listed
-
1 Jul 2018 1 repository listed
-
19 Jun 2018 1 repository listed Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
12 Jun 2018 1 repository listed
-
24 May 2018 1 repository listed
-
3 Apr 2018 1 repository listed
-
1 Apr 2018 1 repository listed
-
Transparency by Design: Closing the Gap Between Performance and Interpretability in Visual Reasoning14 Mar 2018 1 repository listed Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 8 harvested samples)
-
22 Feb 2018 1 repository listed
-
15 Feb 2018 1 repository listed
-
15 Feb 2018 1 repository listed Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
1 Feb 2018 1 repository listed
-
29 Jan 2018 1 repository listed
-
24 Jan 2018 1 repository listed
-
24 Jan 2018 1 repository listed
-
9 Dec 2017 1 repository listed
-
1 Dec 2017 1 repository listed
-
22 Nov 2017 1 repository listed
-
18 Nov 2017 1 repository listed
-
18 Nov 2017 1 repository listed
-
12 Nov 2017 1 repository listed
-
6 Nov 2017 1 repository listed
-
15 Aug 2017 1 repository listed
-
7 Aug 2017 1 repository listed
-
8 Jul 2017 1 repository listed
-
18 May 2017 1 repository listed
-
1 May 2017 1 repository listed
-
1 May 2017 1 repository listed
-
18 Apr 2017 1 repository listed
-
12 Apr 2017 1 repository listed
-
12 Dec 2016 1 repository listed
-
9 Oct 2016 1 repository listed
-
5 Oct 2016 1 repository listed
-
20 Jul 2016 1 repository listed
-
23 Jun 2016 1 repository listed Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
5 Jun 2016 1 repository listed
-
9 May 2016 1 repository listed
-
12 Apr 2016 1 repository listed
-
17 Nov 2015 1 repository listed
-
3 Jun 2015 1 repository listed
-
Barriers in Integrating Medical Visual Question Answering into Radiology Workflows: A Scoping Review and Clinicians' Insights9 Jul 2025 0 repositories listed
-
Evaluating Attribute Confusion in Fashion Text-to-Image Generation9 Jul 2025 0 repositories listed
-
LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation9 Jul 2025 0 repositories listed
-
MagiC: Evaluating Multimodal Cognition Toward Grounded Visual Reasoning9 Jul 2025 0 repositories listed
-
Enhancing Scientific Visual Question Answering through Multimodal Reasoning and Ensemble Modeling8 Jul 2025 0 repositories listed
-
ReLoop: "Seeing Twice and Thinking Backwards" via Closed-loop Training to Mitigate Hallucinations in Multimodal understanding7 Jul 2025 0 repositories listed
-
SMMILE: An Expert-Driven Benchmark for Multimodal Medical In-Context Learning26 Jun 2025 0 repositories listed
-
25 Jun 2025 0 repositories listed
-
Semantic-enhanced Modality-asymmetric Retrieval for Online E-commerce Search25 Jun 2025 0 repositories listed
-
GEMeX-ThinkVG: Towards Thinking with Visual Grounding in Medical VQA via Reinforcement Learning22 Jun 2025 0 repositories listed
-
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations21 Jun 2025 0 repositories listed
-
Can Common VLMs Rival Medical VLMs? Evaluation and Strategic Insights19 Jun 2025 0 repositories listed
-
MEGC2025: Micro-Expression Grand Challenge on Spot Then Recognize and Visual Question Answering18 Jun 2025 0 repositories listed
-
CAPO: Reinforcing Consistent Reasoning in Medical Decision-Making15 Jun 2025 0 repositories listed
-
AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making14 Jun 2025 0 repositories listed
-
A Fast, Reliable, and Secure Programming Language for LLM Agents with Code Actions13 Jun 2025 0 repositories listed
-
MTabVQA: Evaluating Multi-Tabular Reasoning of Language Models in Visual Space13 Jun 2025 0 repositories listed
-
HalLoc: Token-level Localization of Hallucinations for Vision Language Models12 Jun 2025 0 repositories listed
-
Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy11 Jun 2025 0 repositories listed
-
Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning11 Jun 2025 0 repositories listed
-
PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly10 Jun 2025 0 repositories listed
-
Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning8 Jun 2025 0 repositories listed
-
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning8 Jun 2025 0 repositories listed
-
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering7 Jun 2025 0 repositories listed
-
Ontology-based knowledge representation for bone disease diagnosis: a foundation for safe and sustainable medical artificial intelligence systems5 Jun 2025 0 repositories listed
-
TextVidBench: A Benchmark for Long Video Scene Text Understanding5 Jun 2025 0 repositories listed
-
ReXVQA: A Large-scale Visual Question Answering Benchmark for Generalist Chest X-ray Understanding4 Jun 2025 0 repositories listed
-
Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and Transcreation2 Jun 2025 0 repositories listed
-
Learning Sparsity for Effective and Efficient Music Performance Question Answering2 Jun 2025 0 repositories listed
-
Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering1 Jun 2025 0 repositories listed
-
Light as Deception: GPT-driven Natural Relighting Against Vision-Language Pre-training Models30 May 2025 0 repositories listed
-
MedOrch: Medical Diagnosis with Tool-Augmented Reasoning Agents for Flexible Extensibility30 May 2025 0 repositories listed
-
Vision LLMs Are Bad at Hierarchical Visual Understanding, and LLMs Are the Bottleneck30 May 2025 0 repositories listed
-
mRAG: Elucidating the Design Space of Multi-modal Retrieval-Augmented Generation29 May 2025 0 repositories listed
-
NegVQA: Can Vision Language Models Understand Negation?28 May 2025 0 repositories listed
-
Music's Multimodal Complexity in AVQA: Why We Need More than General Multimodal LLMs27 May 2025 0 repositories listed
-
Benchmarking Large Multimodal Models for Ophthalmic Visual Question Answering with OphthalWeChat26 May 2025 0 repositories listed
-
GC-KBVQA: A New Four-Stage Framework for Enhancing Knowledge Based Visual Question Answering Performance25 May 2025 0 repositories listed
-
A Causal Approach to Mitigate Modality Preference Bias in Medical Visual Question Answering22 May 2025 0 repositories listed
-
CT-Agent: A Multimodal-LLM Agent for 3D CT Radiology Question Answering22 May 2025 0 repositories listed
-
Grounding Chest X-Ray Visual Question Answering with Generated Radiology Reports22 May 2025 0 repositories listed
-
Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding22 May 2025 0 repositories listed
-
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation22 May 2025 0 repositories listed
-
Zero-Shot Anomaly Detection in Battery Thermal Images Using Visual Question Answering with Prior Knowledge22 May 2025 0 repositories listed
-
Discovering Pathology Rationale and Token Allocation for Efficient Multimodal Pathology Reasoning21 May 2025 0 repositories listed
-
Human-centered Interactive Learning via MLLMs for Text-to-Image Person Re-identification21 May 2025 0 repositories listed
-
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets21 May 2025 0 repositories listed
-
TinyDrive: Multiscale Visual Question Answering with Selective Token Routing for Autonomous Driving21 May 2025 0 repositories listed
-
Visual Question Answering on Multiple Remote Sensing Image Modalities21 May 2025 0 repositories listed
-
Debating for Better Reasoning: An Unsupervised Multimodal Approach20 May 2025 0 repositories listed
-
Domain Adaptation of VLM for Soccer Video Understanding20 May 2025 0 repositories listed
-
Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models20 May 2025 0 repositories listed
-
Towards Omnidirectional Reasoning with 360-R1: A Dataset, Benchmark, and GRPO-based Method20 May 2025 0 repositories listed
-
Understanding Complexity in VideoQA via Visual Program Generation19 May 2025 0 repositories listed
-
End-to-End Vision Tokenizer Tuning15 May 2025 0 repositories listed
-
Variational Visual Question Answering14 May 2025 0 repositories listed
-
Visually Interpretable Subtask Reasoning for Visual Question Answering12 May 2025 0 repositories listed
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.