Browse State-of-the-Art › Visual Question Answering (VQA) › Papers, page 11
Visual Question Answering (VQA)
Papers archive 2025-07-28
archive papers tagged: 2,167 · with a code link: 1,039 · where Syntology ran a sample: 359 (287 with a run with no instrument failure, 72 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (359 of 2,167 tagged: 287 with a run with no instrument failure, 72 where every run was a failure of Syntology's instrument)
Page 11 of 22: papers 1,001 to 1,100 of 2,167, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
22 Feb 2018 1 repository listed
-
15 Feb 2018 1 repository listed
-
15 Feb 2018 1 repository listed Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
1 Feb 2018 1 repository listed
-
29 Jan 2018 1 repository listed
-
24 Jan 2018 1 repository listed
-
24 Jan 2018 1 repository listed
-
9 Dec 2017 1 repository listed
-
1 Dec 2017 1 repository listed
-
22 Nov 2017 1 repository listed
-
18 Nov 2017 1 repository listed
-
18 Nov 2017 1 repository listed
-
12 Nov 2017 1 repository listed
-
6 Nov 2017 1 repository listed
-
19 Oct 2017 1 repository listed Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
15 Aug 2017 1 repository listed
-
7 Aug 2017 1 repository listed
-
8 Jul 2017 1 repository listed
-
18 May 2017 1 repository listed
-
1 May 2017 1 repository listed
-
1 May 2017 1 repository listed
-
18 Apr 2017 1 repository listed
-
12 Apr 2017 1 repository listed
-
12 Dec 2016 1 repository listed
-
9 Oct 2016 1 repository listed
-
5 Oct 2016 1 repository listed
-
4 Oct 2016 1 repository listed
-
20 Jul 2016 1 repository listed
-
23 Jun 2016 1 repository listed Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
5 Jun 2016 1 repository listed
-
9 May 2016 1 repository listed
-
12 Apr 2016 1 repository listed
-
24 Mar 2016 1 repository listed Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
18 Nov 2015 1 repository listed
-
17 Nov 2015 1 repository listed
-
3 Jun 2015 1 repository listed
-
21 May 2015 1 repository listed
-
9 Jan 2014 1 repository listed
-
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning17 Jul 2025 0 repositories listed
-
MGFFD-VLM: Multi-Granularity Prompt Learning for Face Forgery Detection with VLM16 Jul 2025 0 repositories listed
-
Evaluating Attribute Confusion in Fashion Text-to-Image Generation9 Jul 2025 0 repositories listed
-
LinguaMark: Do Multimodal Models Speak Fairly? A Benchmark-Based Evaluation9 Jul 2025 0 repositories listed
-
Bridging Video Quality Scoring and Justification via Large Multimodal Models26 Jun 2025 0 repositories listed
-
SMMILE: An Expert-Driven Benchmark for Multimodal Medical In-Context Learning26 Jun 2025 0 repositories listed
-
25 Jun 2025 0 repositories listed
-
GEMeX-ThinkVG: Towards Thinking with Visual Grounding in Medical VQA via Reinforcement Learning22 Jun 2025 0 repositories listed
-
Can Common VLMs Rival Medical VLMs? Evaluation and Strategic Insights19 Jun 2025 0 repositories listed
-
How Far Can Off-the-Shelf Multimodal Large Language Models Go in Online Episodic Memory Question Answering?19 Jun 2025 0 repositories listed
-
MEGC2025: Micro-Expression Grand Challenge on Spot Then Recognize and Visual Question Answering18 Jun 2025 0 repositories listed
-
ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM17 Jun 2025 0 repositories listed
-
Connecting phases of matter to the flatness of the loss landscape in analog variational quantum algorithms16 Jun 2025 0 repositories listed
-
CAPO: Reinforcing Consistent Reasoning in Medical Decision-Making15 Jun 2025 0 repositories listed
-
EyeSim-VQA: A Free-Energy-Guided Eye Simulation Framework for Video Quality Assessment13 Jun 2025 0 repositories listed
-
HalLoc: Token-level Localization of Hallucinations for Vision Language Models12 Jun 2025 0 repositories listed
-
Scientists' First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning12 Jun 2025 0 repositories listed
-
Kvasir-VQA-x1: A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy11 Jun 2025 0 repositories listed
-
Provoking Multi-modal Few-Shot LVLM via Exploration-Exploitation In-Context Learning11 Jun 2025 0 repositories listed
-
From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge10 Jun 2025 0 repositories listed
-
PhyBlock: A Progressive Benchmark for Physical Understanding and Planning via 3D Block Assembly10 Jun 2025 0 repositories listed
-
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning8 Jun 2025 0 repositories listed
-
Ontology-based knowledge representation for bone disease diagnosis: a foundation for safe and sustainable medical artificial intelligence systems5 Jun 2025 0 repositories listed
-
ReXVQA: A Large-scale Visual Question Answering Benchmark for Generalist Chest X-ray Understanding4 Jun 2025 0 repositories listed
-
CoRe-MMRAG: Cross-Source Knowledge Reconciliation for Multimodal RAG3 Jun 2025 0 repositories listed
-
Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering1 Jun 2025 0 repositories listed
-
MMedAgent-RL: Optimizing Multi-Agent Collaboration for Multimodal Medical Reasoning31 May 2025 0 repositories listed
-
Proxy-FDA: Proxy-based Feature Distribution Alignment for Fine-tuning Vision Foundation Models without Forgetting30 May 2025 0 repositories listed
-
Vision LLMs Are Bad at Hierarchical Visual Understanding, and LLMs Are the Bottleneck30 May 2025 0 repositories listed
-
A Comprehensive Evaluation of Multi-Modal Large Language Models for Endoscopy Analysis29 May 2025 0 repositories listed
-
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence29 May 2025 0 repositories listed
-
Spoken question answering for visual queries29 May 2025 0 repositories listed
-
NegVQA: Can Vision Language Models Understand Negation?28 May 2025 0 repositories listed
-
Silence is Not Consensus: Disrupting Agreement Bias in Multi-Agent LLMs via Catfish Agent for Clinical Decision Making27 May 2025 0 repositories listed
-
Benchmarking Large Multimodal Models for Ophthalmic Visual Question Answering with OphthalWeChat26 May 2025 0 repositories listed
-
MMIG-Bench: Towards Comprehensive and Explainable Evaluation of Multi-Modal Image Generation Models26 May 2025 0 repositories listed
-
TDVE-Assessor: Benchmarking and Evaluating the Quality of Text-Driven Video Editing with LMMs26 May 2025 0 repositories listed
-
GC-KBVQA: A New Four-Stage Framework for Enhancing Knowledge Based Visual Question Answering Performance25 May 2025 0 repositories listed
-
Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning25 May 2025 0 repositories listed
-
Focus on What Matters: Enhancing Medical Vision-Language Models with Automatic Attention Alignment Tuning24 May 2025 0 repositories listed
-
A Causal Approach to Mitigate Modality Preference Bias in Medical Visual Question Answering22 May 2025 0 repositories listed
-
CT-Agent: A Multimodal-LLM Agent for 3D CT Radiology Question Answering22 May 2025 0 repositories listed
-
Grounding Chest X-Ray Visual Question Answering with Generated Radiology Reports22 May 2025 0 repositories listed
-
MedFrameQA: A Multi-Image Medical VQA Benchmark for Clinical Reasoning22 May 2025 0 repositories listed
-
Steering LVLMs via Sparse Autoencoder for Hallucination Mitigation22 May 2025 0 repositories listed
-
Zero-Shot Anomaly Detection in Battery Thermal Images Using Visual Question Answering with Prior Knowledge22 May 2025 0 repositories listed
-
CP-LLM: Context and Pixel Aware Large Language Model for Video Quality Assessment21 May 2025 0 repositories listed
-
Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning21 May 2025 0 repositories listed
-
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets21 May 2025 0 repositories listed
-
TinyDrive: Multiscale Visual Question Answering with Selective Token Routing for Autonomous Driving21 May 2025 0 repositories listed
-
Visual Question Answering on Multiple Remote Sensing Image Modalities21 May 2025 0 repositories listed
-
Debating for Better Reasoning: An Unsupervised Multimodal Approach20 May 2025 0 repositories listed
-
PlanGPT-VL: Enhancing Urban Planning with Domain-Specific Vision-Language Models20 May 2025 0 repositories listed
-
Toward Effective Reinforcement Learning Fine-Tuning for Medical VQA in Vision-Language Models20 May 2025 0 repositories listed
-
MedSG-Bench: A Benchmark for Medical Image Sequences Grounding17 May 2025 0 repositories listed
-
TinyRS-R1: Compact Multimodal Language Model for Remote Sensing17 May 2025 0 repositories listed
-
Semantically-Aware Game Image Quality Assessment16 May 2025 0 repositories listed
-
Enhancing Multi-Image Question Answering via Submodular Subset Selection15 May 2025 0 repositories listed
-
Variational Visual Question Answering14 May 2025 0 repositories listed
-
10 May 2025 0 repositories listed Syntology 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 5 harvested samples) · 5 pointer-only (licence)
-
Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving9 May 2025 0 repositories listed
Syntology lines on 6 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.