Browse State-of-the-Art › Visual Question Answering › Papers, page 16
Visual Question Answering
Papers archive 2025-07-28
archive papers tagged: 2,177 · with a code link: 1,042 · where Syntology ran a sample: 378 (308 with a run with no instrument failure, 70 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (378 of 2,177 tagged: 308 with a run with no instrument failure, 70 where every run was a failure of Syntology's instrument)
Page 16 of 22: papers 1,501 to 1,600 of 2,177, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Read and Think: An Efficient Step-wise Multimodal Language Model for Document Understanding and Reasoning26 Feb 2024 0 repositories listed
-
25 Feb 2024 0 repositories listed
-
Multimodal Transformer With a Low-Computational-Cost Guarantee23 Feb 2024 0 repositories listed
-
VISREAS: Complex Visual Reasoning with Unanswerable Questions23 Feb 2024 0 repositories listed
-
Exploring the Frontier of Vision-Language Models: A Survey of Current Methodologies and Future Directions20 Feb 2024 0 repositories listed
-
Modality-Aware Integration with Large Language Models for Knowledge-based Visual Question Answering20 Feb 2024 0 repositories listed
-
Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language Models19 Feb 2024 0 repositories listed
-
Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning18 Feb 2024 0 repositories listed
-
PaLM2-VAdapter: Progressively Aligned Language Model Makes a Strong Vision-language Adapter16 Feb 2024 0 repositories listed
-
16 Feb 2024 0 repositories listed Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Prompt-based Personalized Federated Learning for Medical Visual Question Answering15 Feb 2024 0 repositories listed
-
Learning How To Ask: Cycle-Consistency Refines Prompts in Multimodal Foundation Models13 Feb 2024 0 repositories listed
-
Visual Question Answering Instruction: Unlocking Multimodal Large Language Model To Domain-Specific Visual Multitasks13 Feb 2024 0 repositories listed
-
PIVOT: Iterative Visual Prompting Elicits Actionable Knowledge for VLMs12 Feb 2024 0 repositories listed
-
CIC: A Framework for Culturally-Aware Image Captioning8 Feb 2024 0 repositories listed
-
Can Generative AI Support Patients' & Caregivers' Informational Needs? Towards Task-Centric Evaluation Of AI Systems31 Jan 2024 0 repositories listed
-
31 Jan 2024 0 repositories listed
-
LCV2: An Efficient Pretraining-Free Framework for Grounded Visual Question Answering29 Jan 2024 0 repositories listed
-
Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA29 Jan 2024 0 repositories listed
-
Improving Data Augmentation for Robust Visual Question Answering with Effective Curriculum Learning28 Jan 2024 0 repositories listed
-
Free Form Medical Visual Question Answering in Radiology23 Jan 2024 0 repositories listed
-
23 Jan 2024 0 repositories listed
-
22 Jan 2024 0 repositories listed
-
17 Jan 2024 0 repositories listed
-
BOK-VQA: Bilingual outside Knowledge-Based Visual Question Answering via Graph Representation Pretraining12 Jan 2024 0 repositories listed
-
GRAM: Global Reasoning for Multi-Page VQA7 Jan 2024 0 repositories listed
-
4 Jan 2024 0 repositories listed
-
Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers3 Jan 2024 0 repositories listed
-
CoG-DQA: Chain-of-Guiding Learning with Large Language Models for Diagram Question Answering1 Jan 2024 0 repositories listed
-
Mask4Align: Aligned Entity Prompting with Color Masks for Multi-Entity Localization Problems1 Jan 2024 0 repositories listed
-
1 Jan 2024 0 repositories listed
-
Text-Conditioned Generative Model of 3D Strand-based Human Hairstyles1 Jan 2024 0 repositories listed
-
MIVC: Multiple Instance Visual Component for Visual-Language Models28 Dec 2023 0 repositories listed
-
Gemini Pro Defeated by GPT-4V: Evidence from Education27 Dec 2023 0 repositories listed
-
Detection-based Intermediate Supervision for Visual Question Answering26 Dec 2023 0 repositories listed
-
On the Promises and Challenges of Multimodal Foundation Models for Geographical, Environmental, Agricultural, and Urban Planning Applications23 Dec 2023 0 repositories listed
-
Reducing Hallucinations: Enhancing VQA for Flood Disaster Damage Assessment with Visual Contexts21 Dec 2023 0 repositories listed
-
Interactive Visual Task Learning for Robots20 Dec 2023 0 repositories listed
-
Multi-Clue Reasoning with Memory Augmentation for Knowledge-based Visual Question Answering20 Dec 2023 0 repositories listed
-
CLOVA: A Closed-Loop Visual Assistant with Tool Usage and Update18 Dec 2023 0 repositories listed
-
An Evaluation of GPT-4V and Gemini in Online VQA17 Dec 2023 0 repositories listed
-
17 Dec 2023 0 repositories listed
-
BESTMVQA: A Benchmark Evaluation System for Medical Visual Question Answering13 Dec 2023 0 repositories listed
-
Remote Sensing Vision-Language Foundation Models without Annotations via Ground Remote Alignment12 Dec 2023 0 repositories listed
-
Causal-CoG: A Causal-Effect Look at Context Generation for Boosting Multi-modal Language Models9 Dec 2023 0 repositories listed
-
CLAMP: Contrastive LAnguage Model Prompt-tuning4 Dec 2023 0 repositories listed
-
MedXChat: A Unified Multimodal Large Language Model Framework towards CXRs Understanding and Generation4 Dec 2023 0 repositories listed
-
Unleashing the Potential of Large Language Model: Zero-shot VQA for Flood Disaster Scenario4 Dec 2023 0 repositories listed
-
30 Nov 2023 0 repositories listed
-
DreamSync: Aligning Text-to-Image Generation with Image Understanding Feedback29 Nov 2023 0 repositories listed
-
Towards Top-Down Reasoning: An Explainable Multi-Agent Approach for Visual Question Answering29 Nov 2023 0 repositories listed
-
The curse of language biases in remote sensing VQA: the role of spatial attributes, language diversity, and the need for clear evaluation28 Nov 2023 0 repositories listed
-
Lego: Learning to Disentangle and Invert Personalized Concepts Beyond Object Appearance in Text-to-Image Diffusion Models23 Nov 2023 0 repositories listed
-
Asking More Informative Questions for Grounded Retrieval14 Nov 2023 0 repositories listed
-
What Large Language Models Bring to Text-rich VQA?13 Nov 2023 0 repositories listed
-
Visual Commonsense based Heterogeneous Graph Contrastive Learning11 Nov 2023 0 repositories listed
-
CoVLM: Composing Visual Entities and Relationships in Large Language Models Via Communicative Decoding6 Nov 2023 0 repositories listed
-
From Image to Language: A Critical Analysis of Visual Question Answering (VQA) Approaches, Challenges, and Opportunities1 Nov 2023 0 repositories listed
-
VQA-GEN: A Visual Question Answering Benchmark for Domain Generalization1 Nov 2023 0 repositories listed
-
A Systematic Evaluation of GPT-4V's Multimodal Capability for Medical Image Analysis31 Oct 2023 0 repositories listed
-
Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation27 Oct 2023 0 repositories listed
-
25 Oct 2023 0 repositories listed
-
Enhancing Document Information Analysis with Multi-Task Pre-training: A Robust Approach for Information Extraction in Visually-Rich Documents25 Oct 2023 0 repositories listed
-
Exploring Question Decomposition for Zero-Shot VQA25 Oct 2023 0 repositories listed
-
Multimodal Representations for Teacher-Guided Compositional Visual Reasoning24 Oct 2023 0 repositories listed
-
Dataset Bias Mitigation in Multiple-Choice Visual Question Answering and Beyond23 Oct 2023 0 repositories listed
-
20 Oct 2023 0 repositories listed Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
20 Oct 2023 0 repositories listed
-
Enhancing BERT-Based Visual Question Answering through Keyword-Driven Sentence Selection13 Oct 2023 0 repositories listed
-
Exploring Sparse Spatial Relation in Graph Inference for Text-Based VQA13 Oct 2023 0 repositories listed
-
Ziya-Visual: Bilingual Large Vision-Language Model via Multi-Task Instruction Tuning12 Oct 2023 0 repositories listed
-
Improving mitosis detection on histopathology images using large vision-language models11 Oct 2023 0 repositories listed
-
Jaeger: A Concatenation-Based Multi-Transformer VQA Model11 Oct 2023 0 repositories listed
-
Solution for SMART-101 Challenge of ICCV Multi-modal Algorithmic Reasoning Task 202310 Oct 2023 0 repositories listed
-
Causal Reasoning through Two Layers of Cognition for Improving Generalization in Visual Question Answering9 Oct 2023 0 repositories listed
-
Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models9 Oct 2023 0 repositories listed
-
Lightweight In-Context Tuning for Multimodal Unified Models8 Oct 2023 0 repositories listed
-
Improving Automatic VQA Evaluation Using Large Language Models4 Oct 2023 0 repositories listed
-
On the Cognition of Visual Question Answering Models and Human Intelligence: A Comparative Study4 Oct 2023 0 repositories listed
-
SelfGraphVQA: A Self-Supervised Graph Neural Network for Scene-based Question Answering3 Oct 2023 0 repositories listed
-
Human Mobility Question Answering (Vision Paper)2 Oct 2023 0 repositories listed
-
Tackling VQA with Pretrained Foundation Models without Further Training27 Sep 2023 0 repositories listed
-
KOSMOS-2.5: A Multimodal Literate Model20 Sep 2023 0 repositories listed
-
Sentence Attention Blocks for Answer Grounding20 Sep 2023 0 repositories listed
-
Visual Question Answering in the Medical Domain20 Sep 2023 0 repositories listed
-
Syntax Tree Constrained Graph Network for Visual Question Answering17 Sep 2023 0 repositories listed
-
Rank2Tell: A Multimodal Driving Dataset for Joint Importance Ranking and Reasoning12 Sep 2023 0 repositories listed
-
Evaluation and Enhancement of Semantic Grounding in Large Vision-Language Models7 Sep 2023 0 repositories listed
-
Interpretable Visual Question Answering via Reasoning Supervision7 Sep 2023 0 repositories listed
-
Physically Grounded Vision-Language Models for Robotic Manipulation5 Sep 2023 0 repositories listed
-
Expanding Frozen Vision-Language Models without Retraining: Towards Improved Robot Perception31 Aug 2023 0 repositories listed
-
DLIP: Distilling Language-Image Pre-training24 Aug 2023 0 repositories listed
-
EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE23 Aug 2023 0 repositories listed
-
SCULPT: Shape-Conditioned Unpaired Learning of Pose-dependent Clothed and Textured Human Meshes21 Aug 2023 0 repositories listed
-
Generic Attention-model Explainability by Weighted Relevance Accumulation20 Aug 2023 0 repositories listed
-
Towards Grounded Visual Spatial Reasoning in Multi-Modal Vision Language Models18 Aug 2023 0 repositories listed
-
TIJO: Trigger Inversion with Joint Optimization for Defending Multimodal Backdoored Models7 Aug 2023 0 repositories listed
-
ELIXR: Towards a general purpose X-ray artificial intelligence system through alignment of large language models and radiology vision encoders2 Aug 2023 0 repositories listed
-
BARTPhoBEiT: Pre-trained Sequence-to-Sequence and Image Transformers Models for Vietnamese Visual Question Answering28 Jul 2023 0 repositories listed
-
LOIS: Looking Out of Instance Semantics for Visual Question Answering26 Jul 2023 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.