Browse State-of-the-Art › Visual Question Answering (VQA) › Papers, page 12
Visual Question Answering (VQA)
Papers archive 2025-07-28
archive papers tagged: 2,167 · with a code link: 1,039 · where Syntology ran a sample: 359 (287 with a run with no instrument failure, 72 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (359 of 2,167 tagged: 287 with a run with no instrument failure, 72 where every run was a failure of Syntology's instrument)
Page 12 of 22: papers 1,101 to 1,200 of 2,167, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
R^3-VQA: "Read the Room" by Video Social Reasoning7 May 2025 0 repositories listed
-
DiffVQA: Video Quality Assessment Using Diffusion Feature Extractor6 May 2025 0 repositories listed
-
AOR: Anatomical Ontology-Guided Reasoning for Medical Large Multimodal Model in Chest X-Ray Interpretation5 May 2025 0 repositories listed
-
Task-Oriented Semantic Communication in Large Multimodal Models-based Vehicle Networks5 May 2025 0 repositories listed
-
SpatialLLM: A Compound 3D-Informed Design towards Spatially-Intelligent Large Multimodal Models1 May 2025 0 repositories listed
-
Localizing Before Answering: A Hallucination Evaluation Benchmark for Grounded Medical Multimodal LLMs30 Apr 2025 0 repositories listed
-
A Comprehensive Survey of Knowledge-Based Vision Question Answering Systems: The Lifecycle of Knowledge in Visual Reasoning Task24 Apr 2025 0 repositories listed
-
Data-Driven Calibration of Prediction Sets in Large Vision-Language Models Based on Inductive Conformal Prediction24 Apr 2025 0 repositories listed
-
Bridging the Semantic Gaps: Improving Medical VQA Consistency with LLM-Augmented Question Sets16 Apr 2025 0 repositories listed
-
DVLTA-VQA: Decoupled Vision-Language Modeling with Text-Guided Adaptation for Blind Video Quality Assessment16 Apr 2025 0 repositories listed
-
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching16 Apr 2025 0 repositories listed
-
PuzzleBench: A Fully Dynamic Evaluation Framework for Large Multimodal Models on Puzzle Solving15 Apr 2025 0 repositories listed
-
Building Trustworthy Multimodal AI: A Review of Fairness, Transparency, and Ethics in Vision-Language Tasks14 Apr 2025 0 repositories listed
-
MMKB-RAG: A Multi-Modal Knowledge-Based Retrieval-Augmented Generation Framework14 Apr 2025 0 repositories listed
-
NoTeS-Bank: Benchmarking Neural Transcription and Search for Scientific Notes Understanding12 Apr 2025 0 repositories listed
-
PathVLM-R1: A Reinforcement Learning-Driven Reasoning Model for Pathology Visual-Language Tasks12 Apr 2025 0 repositories listed
-
TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs10 Apr 2025 0 repositories listed
-
UniRVQA: A Unified Framework for Retrieval-Augmented Vision Question Answering via Self-Reflective Joint Training5 Apr 2025 0 repositories listed
-
Hierarchical Modeling for Medical Visual Question Answering with Cross-Attention Fusion4 Apr 2025 0 repositories listed
-
QIRL: Boosting Visual Question Answering via Optimized Question-Image Relation Learning4 Apr 2025 0 repositories listed
-
SocialGesture: Delving into Multi-person Gesture Understanding3 Apr 2025 0 repositories listed
-
Reasoning LLMs for User-Aware Multimodal Conversational Agents2 Apr 2025 0 repositories listed
-
MPDrive: Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous Driving1 Apr 2025 0 repositories listed
-
How Well Can Vison-Language Models Understand Humans' Intention? An Open-ended Theory of Mind Question Evaluation Benchmark28 Mar 2025 0 repositories listed
-
Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields26 Mar 2025 0 repositories listed
-
Vision-Amplified Semantic Entropy for Hallucination Detection in Medical Visual Question Answering26 Mar 2025 0 repositories listed
-
LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?25 Mar 2025 0 repositories listed
-
25 Mar 2025 0 repositories listed
-
DiN: Diffusion Model for Robust Medical VQA with Semantic Noisy Labels24 Mar 2025 0 repositories listed
-
MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering24 Mar 2025 0 repositories listed
-
Where is this coming from? Making groundedness count in the evaluation of Document VQA models24 Mar 2025 0 repositories listed
-
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models23 Mar 2025 0 repositories listed
-
A Vision Centric Remote Sensing Benchmark20 Mar 2025 0 repositories listed
-
TruthLens:A Training-Free Paradigm for DeepFake Detection19 Mar 2025 0 repositories listed
-
UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model Evaluation19 Mar 2025 0 repositories listed
-
ChatBEV: A Visual Language Model that Understands BEV Maps18 Mar 2025 0 repositories listed
-
18 Mar 2025 0 repositories listed Syntology 3 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 3 samples that ran constructed an object rather than computing a result (of 3 harvested samples) · 3 pointer-only (licence)
-
GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing16 Mar 2025 0 repositories listed
-
DynRsl-VLM: Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models14 Mar 2025 0 repositories listed
-
Astrea: A MOE-based Visual Understanding Model with Progressive Alignment12 Mar 2025 0 repositories listed
-
SurgicalVLM-Agent: Towards an Interactive AI Co-Pilot for Pituitary Surgery12 Mar 2025 0 repositories listed
-
Bring Remote Sensing Object Detect Into Nature Language Model: Using SFT Method11 Mar 2025 0 repositories listed
-
ComicsPAP: understanding comic strips by picking the correct panel11 Mar 2025 0 repositories listed
-
Seeing and Reasoning with Confidence: Supercharging Multimodal LLMs with an Uncertainty-Aware Agentic Framework11 Mar 2025 0 repositories listed
-
Robusto-1 Dataset: Comparing Humans and VLMs on real out-of-distribution Autonomous Driving VQA from Peru10 Mar 2025 0 repositories listed
-
CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model9 Mar 2025 0 repositories listed
-
Integrating Frequency-Domain Representations with Low-Rank Adaptation in Vision-Language Models8 Mar 2025 0 repositories listed
-
MoEMoE: Question Guided Dense and Scalable Sparse Mixture-of-Expert for Multi-source Multi-modal Answering8 Mar 2025 0 repositories listed
-
SplatTalk: 3D VQA with Gaussian Splatting8 Mar 2025 0 repositories listed
-
4 Mar 2025 0 repositories listed Syntology 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 6 harvested samples) · 6 pointer-only (licence)
-
Enhancing Multi-hop Reasoning in Vision-Language Models via Self-Distillation with Multi-Prompt Ensembling3 Mar 2025 0 repositories listed
-
V²Dial: Unification of Video and Visual Dialog via Multimodal Experts3 Mar 2025 0 repositories listed
-
FunBench: Benchmarking Fundus Reading Skills of MLLMs2 Mar 2025 0 repositories listed
-
ABC: Achieving Better Control of Multimodal Embeddings using VLMs1 Mar 2025 0 repositories listed
-
CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering1 Mar 2025 0 repositories listed
-
Fine-Grained Retrieval-Augmented Generation for Visual Question Answering28 Feb 2025 0 repositories listed
-
ChatReID: Open-ended Interactive Person Retrieval via Hierarchical Progressive Tuning for Vision Language Models27 Feb 2025 0 repositories listed
-
Talking to the brain: Using Large Language Models as Proxies to Model Brain Semantic Representation26 Feb 2025 0 repositories listed
-
FilterRAG: Zero-Shot Informed Retrieval-Augmented Generation to Mitigate Hallucinations in VQA25 Feb 2025 0 repositories listed
-
Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search Engines23 Feb 2025 0 repositories listed
-
Directional Gradient Projection for Robust Fine-Tuning of Foundation Models21 Feb 2025 0 repositories listed
-
Exploring Advanced Techniques for Visual Question Answering: A Comprehensive Comparison20 Feb 2025 0 repositories listed
-
Hardware-Friendly Static Quantization Method for Video Diffusion Transformers20 Feb 2025 0 repositories listed
-
Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning19 Feb 2025 0 repositories listed
-
SafeEraser: Enhancing Safety in Multimodal Large Language Models through Multimodal Machine Unlearning18 Feb 2025 0 repositories listed
-
Multi-Modal Retrieval Augmentation for Open-Ended and Knowledge-Intensive Video Question Answering17 Feb 2025 0 repositories listed
-
14 Feb 2025 0 repositories listed
-
Abduction of Domain Relationships from Data for VQA13 Feb 2025 0 repositories listed
-
EmoAssist: Emotional Assistant for Visual Impairment Community13 Feb 2025 0 repositories listed
-
Visual Graph Question Answering with ASP and LLMs for Language Parsing13 Feb 2025 0 repositories listed
-
Performance Analysis of Traditional VQA Models Under Limited Computational Resources9 Feb 2025 0 repositories listed
-
Hummingbird: High Fidelity Image Generation via Multimodal Context Alignment7 Feb 2025 0 repositories listed
-
Efficient Few-Shot Continual Learning in Vision-Language Models6 Feb 2025 0 repositories listed
-
HD-EPIC: A Highly-Detailed Egocentric Video Dataset6 Feb 2025 0 repositories listed
-
Hypo3D: Exploring Hypothetical Reasoning in 3D2 Feb 2025 0 repositories listed
-
VLM-Assisted Continual learning for Visual Question Answering in Self-Driving2 Feb 2025 0 repositories listed
-
Scaling Large Vision-Language Models for Enhanced Multimodal Comprehension In Biomedical Image Analysis26 Jan 2025 0 repositories listed
-
Scene Understanding Enabled Semantic Communication with Open Channel Coding24 Jan 2025 0 repositories listed
-
Combining Knowledge Graph and LLMs for Enhanced Zero-shot Visual Question Answering22 Jan 2025 0 repositories listed
-
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!18 Jan 2025 0 repositories listed
-
Embodied Scene Understanding for Vision Language Models via MetaVQA15 Jan 2025 0 repositories listed
-
The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering13 Jan 2025 0 repositories listed
-
Video Quality Assessment for Online Processing: From Spatial to Temporal Sampling13 Jan 2025 0 repositories listed
-
Overcoming Language Priors for Visual Question Answering Based on Knowledge Distillation10 Jan 2025 0 repositories listed
-
Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning9 Jan 2025 0 repositories listed
-
Enhancing Financial VQA in Vision Language Models using Intermediate Structured Representations8 Jan 2025 0 repositories listed
-
Visual question answering: from early developments to recent advances -- a survey7 Jan 2025 0 repositories listed
-
Multilevel Semantic-Aware Model for AI-Generated Video Quality Assessment6 Jan 2025 0 repositories listed
-
Accounting for Focus Ambiguity in Visual Questions4 Jan 2025 0 repositories listed
-
Guiding Medical Vision-Language Models with Explicit Visual Prompts: Framework Design and Comprehensive Exploration of Prompt Variations4 Jan 2025 0 repositories listed
-
Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models3 Jan 2025 0 repositories listed
-
MoColl: Agent-Based Specific and General Model Collaboration for Image Captioning3 Jan 2025 0 repositories listed
-
CLIP-UP: CLIP-Based Unanswerable Problem Detection for Visual Question Answering2 Jan 2025 0 repositories listed
-
Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question Answering1 Jan 2025 0 repositories listed
-
F^3OCUS - Federated Finetuning of Vision-Language Foundation Models with Optimal Client Layer Updating Strategy via Multi-objective Meta-Heuristics1 Jan 2025 0 repositories listed
-
JTD-UAV: MLLM-Enhanced Joint Tracking and Description Framework for Anti-UAV Systems1 Jan 2025 0 repositories listed
-
V^2Dial: Unification of Video and Visual Dialog via Multimodal Experts1 Jan 2025 0 repositories listed
-
Probing Visual Language Priors in VLMs31 Dec 2024 0 repositories listed
-
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering30 Dec 2024 0 repositories listed
-
Investigating layer-selective transfer learning of QAOA parameters for Max-Cut problem30 Dec 2024 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.