Browse State-of-the-Art › Visual Question Answering (VQA) › Papers, page 15
Visual Question Answering (VQA)
Papers archive 2025-07-28
archive papers tagged: 2,167 · with a code link: 1,039 · where Syntology ran a sample: 359 (287 with a run with no instrument failure, 72 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (359 of 2,167 tagged: 287 with a run with no instrument failure, 72 where every run was a failure of Syntology's instrument)
Page 15 of 22: papers 1,401 to 1,500 of 2,167, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Visual Question Answering Instruction: Unlocking Multimodal Large Language Model To Domain-Specific Visual Multitasks13 Feb 2024 0 repositories listed
-
CIC: A Framework for Culturally-Aware Image Captioning8 Feb 2024 0 repositories listed
-
Curriculum reinforcement learning for quantum architecture search under hardware errors5 Feb 2024 0 repositories listed
-
Binding Touch to Everything: Learning Unified Multimodal Tactile Representations31 Jan 2024 0 repositories listed
-
LCV2: An Efficient Pretraining-Free Framework for Grounded Visual Question Answering29 Jan 2024 0 repositories listed
-
Muffin or Chihuahua? Challenging Multimodal Large Language Models with Multipanel VQA29 Jan 2024 0 repositories listed
-
Improving Data Augmentation for Robust Visual Question Answering with Effective Curriculum Learning28 Jan 2024 0 repositories listed
-
Free Form Medical Visual Question Answering in Radiology23 Jan 2024 0 repositories listed
-
22 Jan 2024 0 repositories listed
-
Advancing Large Multi-modal Models with Explicit Chain-of-Reasoning and Visual Question Generation18 Jan 2024 0 repositories listed
-
17 Jan 2024 0 repositories listed
-
Video Quality Assessment Based on Swin TransformerV2 and Coarse to Fine Strategy16 Jan 2024 0 repositories listed
-
BOK-VQA: Bilingual outside Knowledge-Based Visual Question Answering via Graph Representation Pretraining12 Jan 2024 0 repositories listed
-
GRAM: Global Reasoning for Multi-Page VQA7 Jan 2024 0 repositories listed
-
DIEM: Decomposition-Integration Enhancing Multimodal Insights1 Jan 2024 0 repositories listed
-
Mask4Align: Aligned Entity Prompting with Color Masks for Multi-Entity Localization Problems1 Jan 2024 0 repositories listed
-
1 Jan 2024 0 repositories listed
-
Text-Conditioned Generative Model of 3D Strand-based Human Hairstyles1 Jan 2024 0 repositories listed
-
Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision Language Audio and Action1 Jan 2024 0 repositories listed
-
Multi-Prompts Learning with Cross-Modal Alignment for Attribute-based Person Re-Identification28 Dec 2023 0 repositories listed
-
Gemini Pro Defeated by GPT-4V: Evidence from Education27 Dec 2023 0 repositories listed
-
Q-Boost: On Visual Quality Assessment Ability of Low-level Multi-Modality Foundation Models23 Dec 2023 0 repositories listed
-
LLM4VG: Large Language Models Evaluation for Video Grounding21 Dec 2023 0 repositories listed
-
Reducing Hallucinations: Enhancing VQA for Flood Disaster Damage Assessment with Visual Contexts21 Dec 2023 0 repositories listed
-
BloomVQA: Assessing Hierarchical Multi-modal Comprehension20 Dec 2023 0 repositories listed
-
Interactive Visual Task Learning for Robots20 Dec 2023 0 repositories listed
-
Multi-Clue Reasoning with Memory Augmentation for Knowledge-based Visual Question Answering20 Dec 2023 0 repositories listed
-
Full-reference Video Quality Assessment for User Generated Content Transcoding19 Dec 2023 0 repositories listed
-
An Evaluation of GPT-4V and Gemini in Online VQA17 Dec 2023 0 repositories listed
-
RankDVQA-mini: Knowledge Distillation-Driven Deep Video Quality Assessment14 Dec 2023 0 repositories listed
-
BESTMVQA: A Benchmark Evaluation System for Medical Visual Question Answering13 Dec 2023 0 repositories listed
-
Causal-CoG: A Causal-Effect Look at Context Generation for Boosting Multi-modal Language Models9 Dec 2023 0 repositories listed
-
8 Dec 2023 0 repositories listed
-
On the Robustness of Large Multimodal Models Against Image Adversarial Attacks6 Dec 2023 0 repositories listed
-
Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models5 Dec 2023 0 repositories listed
-
MedXChat: A Unified Multimodal Large Language Model Framework towards CXRs Understanding and Generation4 Dec 2023 0 repositories listed
-
Unleashing the Potential of Large Language Model: Zero-shot VQA for Flood Disaster Scenario4 Dec 2023 0 repositories listed
-
1 Dec 2023 0 repositories listed
-
1 Dec 2023 0 repositories listed
-
Towards Top-Down Reasoning: An Explainable Multi-Agent Approach for Visual Question Answering29 Nov 2023 0 repositories listed
-
The curse of language biases in remote sensing VQA: the role of spatial attributes, language diversity, and the need for clear evaluation28 Nov 2023 0 repositories listed
-
From Wrong To Right: A Recursive Approach Towards Vision-Language Explanation21 Nov 2023 0 repositories listed
-
KNVQA: A Benchmark for evaluation knowledge-based VQA21 Nov 2023 0 repositories listed
-
Understanding and Mitigating Classification Errors Through Interpretable Token Patterns18 Nov 2023 0 repositories listed
-
Multiple-Question Multiple-Answer Text-VQA15 Nov 2023 0 repositories listed
-
Asking More Informative Questions for Grounded Retrieval14 Nov 2023 0 repositories listed
-
CLiF-VQA: Enhancing Video Quality Assessment by Incorporating High-Level Semantic Information related to Human Feelings13 Nov 2023 0 repositories listed
-
What Large Language Models Bring to Text-rich VQA?13 Nov 2023 0 repositories listed
-
Visual Commonsense based Heterogeneous Graph Contrastive Learning11 Nov 2023 0 repositories listed
-
Improving Vision-and-Language Reasoning via Spatial Relations Modeling9 Nov 2023 0 repositories listed
-
From Image to Language: A Critical Analysis of Visual Question Answering (VQA) Approaches, Challenges, and Opportunities1 Nov 2023 0 repositories listed
-
VQA-GEN: A Visual Question Answering Benchmark for Domain Generalization1 Nov 2023 0 repositories listed
-
A Systematic Evaluation of GPT-4V's Multimodal Capability for Medical Image Analysis31 Oct 2023 0 repositories listed
-
Davidsonian Scene Graph: Improving Reliability in Fine-grained Evaluation for Text-to-Image Generation27 Oct 2023 0 repositories listed
-
Exploring Question Decomposition for Zero-Shot VQA25 Oct 2023 0 repositories listed
-
Geometry-Aware Video Quality Assessment for Dynamic Digital Human24 Oct 2023 0 repositories listed
-
20 Oct 2023 0 repositories listed Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Exploring Sparse Spatial Relation in Graph Inference for Text-Based VQA13 Oct 2023 0 repositories listed
-
Improving mitosis detection on histopathology images using large vision-language models11 Oct 2023 0 repositories listed
-
Jaeger: A Concatenation-Based Multi-Transformer VQA Model11 Oct 2023 0 repositories listed
-
Off-Policy Evaluation for Human Feedback11 Oct 2023 0 repositories listed
-
How (not) to ensemble LVLMs for VQA10 Oct 2023 0 repositories listed
-
Causal Reasoning through Two Layers of Cognition for Improving Generalization in Visual Question Answering9 Oct 2023 0 repositories listed
-
Negative Object Presence Evaluation (NOPE) to Measure Object Hallucination in Vision-Language Models9 Oct 2023 0 repositories listed
-
Improving Automatic VQA Evaluation Using Large Language Models4 Oct 2023 0 repositories listed
-
On the Cognition of Visual Question Answering Models and Human Intelligence: A Comparative Study4 Oct 2023 0 repositories listed
-
SelfGraphVQA: A Self-Supervised Graph Neural Network for Scene-based Question Answering3 Oct 2023 0 repositories listed
-
Tackling VQA with Pretrained Foundation Models without Further Training27 Sep 2023 0 repositories listed
-
Sentence Attention Blocks for Answer Grounding20 Sep 2023 0 repositories listed
-
Visual Question Answering in the Medical Domain20 Sep 2023 0 repositories listed
-
Syntax Tree Constrained Graph Network for Visual Question Answering17 Sep 2023 0 repositories listed
-
Interpretable Visual Question Answering via Reasoning Supervision7 Sep 2023 0 repositories listed
-
S3C: Semi-Supervised VQA Natural Language Explanation via Self-Critical Learning5 Sep 2023 0 repositories listed
-
Distraction-free Embeddings for Robust VQA31 Aug 2023 0 repositories listed
-
Ada-DQA: Adaptive Diverse Quality-aware Feature Acquisition for Video Quality Assessment1 Aug 2023 0 repositories listed
-
Making the V in Text-VQA Matter1 Aug 2023 0 repositories listed
-
Bridging the Gap: Exploring the Capabilities of Bridge-Architectures for Complex Visual Reasoning Tasks31 Jul 2023 0 repositories listed
-
Capturing Co-existing Distortions in User-Generated Content for No-reference Video Quality Assessment31 Jul 2023 0 repositories listed
-
Workshop on Document Intelligence Understanding31 Jul 2023 0 repositories listed
-
BARTPhoBEiT: Pre-trained Sequence-to-Sequence and Image Transformers Models for Vietnamese Visual Question Answering28 Jul 2023 0 repositories listed
-
LOIS: Looking Out of Instance Semantics for Visual Question Answering26 Jul 2023 0 repositories listed
-
Robust Visual Question Answering: Datasets, Methods, and Future Challenges21 Jul 2023 0 repositories listed
-
A reinforcement learning approach for VQA validation: an application to diabetic macular edema grading19 Jul 2023 0 repositories listed
-
NTIRE 2023 Quality Assessment of Video Enhancement Challenge19 Jul 2023 0 repositories listed
-
Generative Visual Question Answering18 Jul 2023 0 repositories listed
-
Let's ViCE! Mimicking Human Cognitive Behavior in Image Generation Evaluation18 Jul 2023 0 repositories listed
-
Divide, Evaluate, and Refine: Evaluating and Improving Text-to-Image Alignment with Iterative VQA Feedback10 Jul 2023 0 repositories listed
-
UIT-Saviors at MEDVQA-GI 2023: Improving Multimodal Learning with Image Enhancement for Gastrointestinal Visual Question Answering6 Jul 2023 0 repositories listed
-
DoReMi: Grounding Language Model by Detecting and Recovering from Plan-Execution Misalignment1 Jul 2023 0 repositories listed
-
Deep Equilibrium Multimodal Fusion29 Jun 2023 0 repositories listed
-
Visual Question Answering in Remote Sensing with Cross-Attention and Multimodal Information Bottleneck25 Jun 2023 0 repositories listed
-
AVIS: Autonomous Visual Information Seeking with Large Language Model Agent13 Jun 2023 0 repositories listed
-
Visual Question Answering (VQA) on Images with Superimposed Text13 Jun 2023 0 repositories listed
-
Weakly Supervised Visual Question Answer Generation11 Jun 2023 0 repositories listed
-
Knowledge Detection by Relevant Question and Image Attributes in Visual Question Answering8 Jun 2023 0 repositories listed
-
Multi-CLIP: Contrastive Vision-Language Pre-training for Question Answering tasks in 3D Scenes4 Jun 2023 0 repositories listed
-
MetaVL: Transferring In-Context Learning Ability From Language Models to Vision-Language Models2 Jun 2023 0 repositories listed
-
Evaluating the Capabilities of Multi-modal Reasoning Models with Synthetic Task Data1 Jun 2023 0 repositories listed
-
LiT-4-RSVQA: Lightweight Transformer-based Visual Question Answering in Remote Sensing1 Jun 2023 0 repositories listed
-
Overcoming Language Bias in Remote Sensing Visual Question Answering via Adversarial Training1 Jun 2023 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.