Browse State-of-the-Art › Visual Question Answering › Papers, page 14
Visual Question Answering
Papers archive 2025-07-28
archive papers tagged: 2,177 · with a code link: 1,042 · where Syntology ran a sample: 378 (308 with a run with no instrument failure, 70 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (378 of 2,177 tagged: 308 with a run with no instrument failure, 70 where every run was a failure of Syntology's instrument)
Page 14 of 22: papers 1,301 to 1,400 of 2,177, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Interpretable Bilingual Multimodal Large Language Model for Diverse Biomedical Tasks24 Oct 2024 0 repositories listed
-
Which Client is Reliable?: A Reliable and Personalized Prompt-based Federated Learning for Medical Image Question Answering23 Oct 2024 0 repositories listed
-
Order Matters: Exploring Order Sensitivity in Multimodal Large Language Models22 Oct 2024 0 repositories listed
-
Visual Question Answering in Ophthalmology: A Progressive and Practical Perspective22 Oct 2024 0 repositories listed
-
Object-Centric Temporal Consistency via Conditional Autoregressive Inductive Biases21 Oct 2024 0 repositories listed
-
ChitroJera: A Regionally Relevant Visual Question Answering Dataset for Bangla19 Oct 2024 0 repositories listed
-
LLaVA-Ultra: Large Chinese Language and Vision Assistant for Ultrasound19 Oct 2024 0 repositories listed
-
E3D-GPT: Enhanced 3D Visual Foundation for Medical Vision-Language Model18 Oct 2024 0 repositories listed
-
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples18 Oct 2024 0 repositories listed
-
Zero-shot Action Localization via the Confidence of Large Vision-Language Models18 Oct 2024 0 repositories listed
-
17 Oct 2024 0 repositories listed
-
17 Oct 2024 0 repositories listed
-
17 Oct 2024 0 repositories listed
-
RescueADI: Adaptive Disaster Interpretation in Remote Sensing Images with Autonomous Agents17 Oct 2024 0 repositories listed
-
16 Oct 2024 0 repositories listed
-
OMCAT: Omni Context Aware Transformer15 Oct 2024 0 repositories listed
-
Eliminating the Language Bias for Visual Question Answering with fine-grained Causal Intervention14 Oct 2024 0 repositories listed
-
14 Oct 2024 0 repositories listed
-
13 Oct 2024 0 repositories listed
-
Surgical-LLaVA: Toward Surgical Scenario Understanding via Large Language and Vision Models13 Oct 2024 0 repositories listed
-
12 Oct 2024 0 repositories listed Syntology 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
ViT3D Alignment of LLaMA3: 3D Medical Image Report Generation11 Oct 2024 0 repositories listed
-
Emerging Pixel Grounding in Large Multimodal Models Without Grounding Supervision10 Oct 2024 0 repositories listed
-
10 Oct 2024 0 repositories listed
-
PAR: Prompt-Aware Token Reduction Method for Efficient Large Multimodal Models9 Oct 2024 0 repositories listed
-
Beyond Captioning: Task-Specific Prompting for Improved VLM Performance in Mathematical Reasoning8 Oct 2024 0 repositories listed
-
MM-R³: On (In-)Consistency of Multi-modal Large Language Models (MLLMs)7 Oct 2024 0 repositories listed
-
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks7 Oct 2024 0 repositories listed
-
5 Oct 2024 0 repositories listed
-
Backdooring Vision-Language Models with Out-Of-Distribution Data2 Oct 2024 0 repositories listed
-
Why context matters in VQA and Reasoning: Semantic interventions for VLM input modalities2 Oct 2024 0 repositories listed
-
FMBench: Benchmarking Fairness in Multimodal Large Language Models on Medical Tasks1 Oct 2024 0 repositories listed
-
30 Sep 2024 0 repositories listed
-
3D-CT-GPT: Generating 3D Radiology Reports through Integration of Large Vision-Language Models28 Sep 2024 0 repositories listed
-
TrojVLM: Backdoor Attack Against Vision Language Models28 Sep 2024 0 repositories listed
-
Charting the Future: Using Chart Question-Answering for Scalable Evaluation of LLM-Driven Data Visualizations27 Sep 2024 0 repositories listed
-
Enhancing Explainability in Multimodal Large Language Models Using Ontological Context27 Sep 2024 0 repositories listed
-
DARE: Diverse Visual Question Answering with Robustness Evaluation26 Sep 2024 0 repositories listed
-
Robotic Environmental State Recognition with Pre-Trained Vision-Language Models and Black-Box Optimization26 Sep 2024 0 repositories listed
-
ZALM3: Zero-Shot Enhancement of Vision-Language Alignment via In-Context Information in Multi-Turn Multimodal Medical Dialogue26 Sep 2024 0 repositories listed
-
Can CLIP Count Stars? An Empirical Study on Quantity Bias in CLIP23 Sep 2024 0 repositories listed
-
Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation23 Sep 2024 0 repositories listed
-
@Bench: Benchmarking Vision-Language Models for Human-centered Assistive Technology21 Sep 2024 0 repositories listed
-
Vision Language Models Can Parse Floor Plan Maps19 Sep 2024 0 repositories listed
-
OneEncoder: A Lightweight Framework for Progressive Alignment of Modalities17 Sep 2024 0 repositories listed
-
Sparks of Artificial General Intelligence(AGI) in Semiconductor Material Science: Early Explorations into the Next Frontier of Generative AI-Assisted Electron Micrograph Analysis17 Sep 2024 0 repositories listed
-
Explore the Hallucination on Low-level Perception for MLLMs15 Sep 2024 0 repositories listed
-
NEVLP: Noise-Robust Framework for Efficient Vision-Language Pre-training15 Sep 2024 0 repositories listed
-
Learning to Compress Contexts for Efficient Knowledge-based Visual Question Answering11 Sep 2024 0 repositories listed
-
Securing Vision-Language Models with a Robust Encoder Against Jailbreak and Adversarial Attacks11 Sep 2024 0 repositories listed
-
Mitigating Hallucination in Visual-Language Models via Re-Balancing Contrastive Decoding10 Sep 2024 0 repositories listed
-
VisScience: An Extensive Benchmark for Evaluating K12 Educational Multi-modal Scientific Reasoning10 Sep 2024 0 repositories listed
-
Breaking Neural Network Scaling Laws with Modularity9 Sep 2024 0 repositories listed
-
7 Sep 2024 0 repositories listed
-
OccLLaMA: An Occupancy-Language-Action Generative World Model for Autonomous Driving5 Sep 2024 0 repositories listed
-
MOSMOS: Multi-organ segmentation facilitated by medical report supervision4 Sep 2024 0 repositories listed
-
Blocks as Probes: Dissecting Categorization Ability of Large Multimodal Models3 Sep 2024 0 repositories listed
-
Look, Learn and Leverage (L³): Mitigating Visual-Domain Shift and Discovering Intrinsic Relations via Symbolic Alignment30 Aug 2024 0 repositories listed
-
Retrieval-Augmented Natural Language Reasoning for Explainable Visual Question Answering30 Aug 2024 0 repositories listed
-
M4CXR: Exploring Multi-task Potentials of Multi-modal Large Language Models for Chest X-ray Interpretation29 Aug 2024 0 repositories listed
-
Can SAR improve RSVQA performance?28 Aug 2024 0 repositories listed
-
Can Visual Language Models Replace OCR-Based Visual Question Answering Pipelines in Production? A Case Study in Retail28 Aug 2024 0 repositories listed
-
Multi-Modal Instruction-Tuning Small-Scale Language-and-Vision Assistant for Semiconductor Electron Micrograph Analysis27 Aug 2024 0 repositories listed
-
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis27 Aug 2024 0 repositories listed
-
Towards Human-Level Understanding of Complex Process Engineering Schematics: A Pedagogical, Introspective Multi-Agent Framework for Open-Domain Question Answering24 Aug 2024 0 repositories listed
-
Foundational Model for Electron Micrograph Analysis: Instruction-Tuning Small-Scale Language-and-Vision Assistant for Enterprise Adoption23 Aug 2024 0 repositories listed
-
22 Aug 2024 0 repositories listed
-
21 Aug 2024 0 repositories listed
-
Swarm Intelligence in Geo-Localization: A Multi-Agent Large Vision-Language Model Collaborative Framework21 Aug 2024 0 repositories listed
-
Beyond the Hype: A dispassionate look at vision-language models in medical scenario16 Aug 2024 0 repositories listed
-
Enhancing Visual Question Answering through Ranking-Based Hybrid Training and Multimodal Fusion14 Aug 2024 0 repositories listed
-
13 Aug 2024 0 repositories listed
-
Revisiting Multi-Modal LLM Evaluation9 Aug 2024 0 repositories listed
-
8 Aug 2024 0 repositories listed
-
Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble Exploitation7 Aug 2024 0 repositories listed
-
MMPKUBase: A Comprehensive and High-quality Chinese Multi-modal Knowledge Graph3 Aug 2024 0 repositories listed
-
Prompting Medical Large Vision-Language Models to Diagnose Pathologies by Visual Question Answering31 Jul 2024 0 repositories listed
-
SimpleLLM4AD: An End-to-End Vision-Language Model with Graph Visual Question Answering for Autonomous Driving31 Jul 2024 0 repositories listed
-
Pyramid Coder: Hierarchical Code Generator for Compositional Visual Question Answering30 Jul 2024 0 repositories listed
-
Take A Step Back: Rethinking the Two Stages in Visual Reasoning29 Jul 2024 0 repositories listed
-
VolDoGer: LLM-assisted Datasets for Domain Generalization in Vision-Language Tasks29 Jul 2024 0 repositories listed
-
AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question Answering28 Jul 2024 0 repositories listed
-
24 Jul 2024 0 repositories listed
-
Imperfect Vision Encoders: Efficient and Robust Tuning for Vision-Language Models23 Jul 2024 0 repositories listed
-
Exploring the Effectiveness of Object-Centric Representations in Visual Question Answering: Comparative Insights with Foundation Models22 Jul 2024 0 repositories listed
-
EchoSight: Advancing Visual-Language Models with Wiki Knowledge17 Jul 2024 0 repositories listed
-
Multimodal Reranking for Knowledge-Intensive Visual Question Answering17 Jul 2024 0 repositories listed
-
TM-PATHVQA:90000+ Textless Multilingual Questions for Medical Visual Question Answering16 Jul 2024 0 repositories listed
-
Benchmarking Vision Language Models for Cultural Understanding15 Jul 2024 0 repositories listed
-
Extracting Training Data from Document-Based VQA Models11 Jul 2024 0 repositories listed
-
Segmentation-guided Attention for Visual Question Answering from Remote Sensing Images11 Jul 2024 0 repositories listed
-
VQA-Diff: Exploiting VQA and Diffusion for Zero-Shot Image-to-3D Vehicle Asset Generation in Autonomous Driving9 Jul 2024 0 repositories listed
-
5 Jul 2024 0 repositories listed
-
Second Place Solution of WSDM2023 Toloka Visual Question Answering Challenge5 Jul 2024 0 repositories listed
-
BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs3 Jul 2024 0 repositories listed
-
MindBench: A Comprehensive Benchmark for Mind Map Structure Recognition and Analysis3 Jul 2024 0 repositories listed
-
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness2 Jul 2024 0 repositories listed
-
Assistive Image Annotation Systems with Deep Learning and Natural Language Capabilities: A Review28 Jun 2024 0 repositories listed
-
Disentangling Knowledge-based and Visual Reasoning by Question Decomposition in KB-VQA27 Jun 2024 0 repositories listed
-
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts27 Jun 2024 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.