Browse State-of-the-Art › Visual Question Answering › Papers, page 12
Visual Question Answering
Papers archive 2025-07-28
archive papers tagged: 2,177 · with a code link: 1,042 · where Syntology ran a sample: 378 (308 with a run with no instrument failure, 70 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (378 of 2,177 tagged: 308 with a run with no instrument failure, 70 where every run was a failure of Syntology's instrument)
Page 12 of 22: papers 1,101 to 1,200 of 2,177, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Multi-Modal Explainable Medical AI Assistant for Trustworthy Human-AI Collaboration11 May 2025 0 repositories listed
-
10 May 2025 0 repositories listed Syntology 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 5 harvested samples) · 5 pointer-only (licence)
-
Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving9 May 2025 0 repositories listed
-
SITE: towards Spatial Intelligence Thorough Evaluation8 May 2025 0 repositories listed
-
Structure Causal Models and LLMs Integration in Medical Visual Question Answering5 May 2025 0 repositories listed
-
Task-Oriented Semantic Communication in Large Multimodal Models-based Vehicle Networks5 May 2025 0 repositories listed
-
Compositional Image-Text Matching and Retrieval by Grounding Entities4 May 2025 0 repositories listed
-
Adaptive Token Boundaries: Integrating Human Chunking Mechanisms into Multimodal LLMs3 May 2025 0 repositories listed
-
Knowledge-Augmented Language Models Interpreting Structured Chest X-Ray Findings3 May 2025 0 repositories listed
-
Grounding Task Assistance with Multimodal Cues from a Single Demonstration2 May 2025 0 repositories listed
-
Transferable Adversarial Attacks on Black-Box Vision-Language Models2 May 2025 0 repositories listed
-
Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding30 Apr 2025 0 repositories listed
-
LMME3DHF: Benchmarking and Evaluating Multimodal 3D Human Face Generation with LMMs29 Apr 2025 0 repositories listed
-
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning28 Apr 2025 0 repositories listed
-
Data-Driven Calibration of Prediction Sets in Large Vision-Language Models Based on Inductive Conformal Prediction24 Apr 2025 0 repositories listed
-
TraveLLaMA: Facilitating Multi-modal Large Language Models to Understand Urban Scenes and Provide Travel Assistance23 Apr 2025 0 repositories listed
-
Neglected Risks: The Disturbing Reality of Children's Images in Datasets and the Urgent Call for Accountability20 Apr 2025 0 repositories listed
-
Hadamard product in deep learning: Introduction, Advances and Challenges17 Apr 2025 0 repositories listed
-
Bridging the Semantic Gaps: Improving Medical VQA Consistency with LLM-Augmented Question Sets16 Apr 2025 0 repositories listed
-
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching16 Apr 2025 0 repositories listed
-
LVLM_CSP: Accelerating Large Vision Language Models via Clustering, Scattering, and Pruning for Reasoning Segmentation15 Apr 2025 0 repositories listed
-
Building Trustworthy Multimodal AI: A Review of Fairness, Transparency, and Ethics in Vision-Language Tasks14 Apr 2025 0 repositories listed
-
MMKB-RAG: A Multi-Modal Knowledge-Based Retrieval-Augmented Generation Framework14 Apr 2025 0 repositories listed
-
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents14 Apr 2025 0 repositories listed
-
NoTeS-Bank: Benchmarking Neural Transcription and Search for Scientific Notes Understanding12 Apr 2025 0 repositories listed
-
AstroLLaVA: towards the unification of astronomical data and natural language11 Apr 2025 0 repositories listed
-
Beyond the Frame: Generating 360° Panoramic Videos from Perspective Videos10 Apr 2025 0 repositories listed
-
Data Metabolism: An Efficient Data Design Schema For Vision Language Model10 Apr 2025 0 repositories listed
-
TokenFocus-VQA: Enhancing Text-to-Image Alignment with Position-Aware Focus and Multi-Perspective Aggregations on LVLMs10 Apr 2025 0 repositories listed
-
RS-RAG: Bridging Remote Sensing Imagery and Comprehensive Knowledge with a Multi-Modal Dataset and Retrieval-Augmented Generation Model7 Apr 2025 0 repositories listed
-
Hierarchical Modeling for Medical Visual Question Answering with Cross-Attention Fusion4 Apr 2025 0 repositories listed
-
QIRL: Boosting Visual Question Answering via Optimized Question-Image Relation Learning4 Apr 2025 0 repositories listed
-
SocialGesture: Delving into Multi-person Gesture Understanding3 Apr 2025 0 repositories listed
-
MPDrive: Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous Driving1 Apr 2025 0 repositories listed
-
SViQA: A Unified Speech-Vision Multimodal Model for Textless Visual Question Answering1 Apr 2025 0 repositories listed
-
How Well Can Vison-Language Models Understand Humans' Intention? An Open-ended Theory of Mind Question Evaluation Benchmark28 Mar 2025 0 repositories listed
-
CTRL-O: Language-Controllable Object-Centric Visual Representation Learning27 Mar 2025 0 repositories listed
-
JEEM: Vision-Language Understanding in Four Arabic Dialects27 Mar 2025 0 repositories listed
-
Feature4X: Bridging Any Monocular Video to 4D Agentic AI with Versatile Gaussian Feature Fields26 Mar 2025 0 repositories listed
-
Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs26 Mar 2025 0 repositories listed
-
Mitigating Low-Level Visual Hallucinations Requires Self-Awareness: Database, Model and Training Strategy26 Mar 2025 0 repositories listed
-
Vision-Amplified Semantic Entropy for Hallucination Detection in Medical Visual Question Answering26 Mar 2025 0 repositories listed
-
Improved Alignment of Modalities in Large Vision Language Models25 Mar 2025 0 repositories listed
-
LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?25 Mar 2025 0 repositories listed
-
25 Mar 2025 0 repositories listed
-
DiN: Diffusion Model for Robust Medical VQA with Semantic Noisy Labels24 Mar 2025 0 repositories listed
-
MAGIC-VQA: Multimodal And Grounded Inference with Commonsense Knowledge for Visual Question Answering24 Mar 2025 0 repositories listed
-
Where is this coming from? Making groundedness count in the evaluation of Document VQA models24 Mar 2025 0 repositories listed
-
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models23 Mar 2025 0 repositories listed
-
A Vision Centric Remote Sensing Benchmark20 Mar 2025 0 repositories listed
-
EfficientLLaVA:Generalizable Auto-Pruning for Large Vision-language Models19 Mar 2025 0 repositories listed
-
GraspCorrect: Robotic Grasp Correction via Vision-Language Model-Guided Feedback19 Mar 2025 0 repositories listed
-
TruthLens:A Training-Free Paradigm for DeepFake Detection19 Mar 2025 0 repositories listed
-
UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model Evaluation19 Mar 2025 0 repositories listed
-
18 Mar 2025 0 repositories listed Syntology 3 ran (of which 3 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 3 samples that ran constructed an object rather than computing a result (of 3 harvested samples) · 3 pointer-only (licence)
-
From Head to Tail: Towards Balanced Representation in Large Vision-Language Models through Adaptive Data Calibration17 Mar 2025 0 repositories listed
-
Task-Oriented Feature Compression for Multimodal Understanding via Device-Edge Co-Inference17 Mar 2025 0 repositories listed
-
GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing16 Mar 2025 0 repositories listed
-
PEBench: A Fictitious Dataset to Benchmark Machine Unlearning for Multimodal Large Language Models16 Mar 2025 0 repositories listed
-
DynRsl-VLM: Enhancing Autonomous Driving Perception with Dynamic Resolution Vision-Language Models14 Mar 2025 0 repositories listed
-
On the Limitations of Vision-Language Models in Understanding Image Transforms12 Mar 2025 0 repositories listed
-
SurgicalVLM-Agent: Towards an Interactive AI Co-Pilot for Pituitary Surgery12 Mar 2025 0 repositories listed
-
Seeing and Reasoning with Confidence: Supercharging Multimodal LLMs with an Uncertainty-Aware Agentic Framework11 Mar 2025 0 repositories listed
-
From Text to Visuals: Using LLMs to Generate Math Diagrams with Vector Graphics10 Mar 2025 0 repositories listed
-
Robusto-1 Dataset: Comparing Humans and VLMs on real out-of-distribution Autonomous Driving VQA from Peru10 Mar 2025 0 repositories listed
-
TI-JEPA: An Innovative Energy-based Joint Embedding Strategy for Text-Image Multimodal Systems9 Mar 2025 0 repositories listed
-
Integrating Frequency-Domain Representations with Low-Rank Adaptation in Vision-Language Models8 Mar 2025 0 repositories listed
-
MoEMoE: Question Guided Dense and Scalable Sparse Mixture-of-Expert for Multi-source Multi-modal Answering8 Mar 2025 0 repositories listed
-
SplatTalk: 3D VQA with Gaussian Splatting8 Mar 2025 0 repositories listed
-
Enhancing SAM with Efficient Prompting and Preference Optimization for Semi-supervised Medical Image Segmentation6 Mar 2025 0 repositories listed
-
OWLViz: An Open-World Benchmark for Visual Question Answering4 Mar 2025 0 repositories listed
-
FunBench: Benchmarking Fundus Reading Skills of MLLMs2 Mar 2025 0 repositories listed
-
CL-MoE: Enhancing Multimodal Large Language Model with Dual Momentum Mixture-of-Experts for Continual Visual Question Answering1 Mar 2025 0 repositories listed
-
Fine-Grained Retrieval-Augmented Generation for Visual Question Answering28 Feb 2025 0 repositories listed
-
Can Large Language Models Unveil the Mysteries? An Exploration of Their Ability to Unlock Information in Complex Scenarios27 Feb 2025 0 repositories listed
-
MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning26 Feb 2025 0 repositories listed
-
Talking to the brain: Using Large Language Models as Proxies to Model Brain Semantic Representation26 Feb 2025 0 repositories listed
-
Detecting Knowledge Boundary of Vision Large Language Models by Sampling-Based Inference25 Feb 2025 0 repositories listed
-
FilterRAG: Zero-Shot Informed Retrieval-Augmented Generation to Mitigate Hallucinations in VQA25 Feb 2025 0 repositories listed
-
All-in-one: Understanding and Generation in Multimodal Reasoning with the MAIA Benchmark24 Feb 2025 0 repositories listed
-
Retrieval-Augmented Visual Question Answering via Built-in Autoregressive Search Engines23 Feb 2025 0 repositories listed
-
Tracking the Copyright of Large Vision-Language Models through Parameter Learning Adversarial Images23 Feb 2025 0 repositories listed
-
Directional Gradient Projection for Robust Fine-Tuning of Foundation Models21 Feb 2025 0 repositories listed
-
TransMamba: Fast Universal Architecture Adaption from Transformers to Mamba21 Feb 2025 0 repositories listed
-
Exploring Advanced Techniques for Visual Question Answering: A Comprehensive Comparison20 Feb 2025 0 repositories listed
-
Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning19 Feb 2025 0 repositories listed
-
SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models18 Feb 2025 0 repositories listed
-
"See the World, Discover Knowledge": A Chinese Factuality Evaluation for Large Vision Language Models17 Feb 2025 0 repositories listed
-
Abduction of Domain Relationships from Data for VQA13 Feb 2025 0 repositories listed
-
EmoAssist: Emotional Assistant for Visual Impairment Community13 Feb 2025 0 repositories listed
-
Visual Graph Question Answering with ASP and LLMs for Language Parsing13 Feb 2025 0 repositories listed
-
Vision-Language Models for Edge Networks: A Comprehensive Survey11 Feb 2025 0 repositories listed
-
Performance Analysis of Traditional VQA Models Under Limited Computational Resources9 Feb 2025 0 repositories listed
-
Hummingbird: High Fidelity Image Generation via Multimodal Context Alignment7 Feb 2025 0 repositories listed
-
Efficient Few-Shot Continual Learning in Vision-Language Models6 Feb 2025 0 repositories listed
-
Exploring Spatial Language Grounding Through Referring Expressions4 Feb 2025 0 repositories listed
-
Hypo3D: Exploring Hypothetical Reasoning in 3D2 Feb 2025 0 repositories listed
-
VLM-Assisted Continual learning for Visual Question Answering in Self-Driving2 Feb 2025 0 repositories listed
-
Anatomy Might Be All You Need: Forecasting What to Do During Surgery29 Jan 2025 0 repositories listed
-
Scaling Large Vision-Language Models for Enhanced Multimodal Comprehension In Biomedical Image Analysis26 Jan 2025 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.