Browse State-of-the-Art › Visual Question Answering › Papers, page 17
Visual Question Answering
Papers archive 2025-07-28
archive papers tagged: 2,177 · with a code link: 1,042 · where Syntology ran a sample: 378 (308 with a run with no instrument failure, 70 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (378 of 2,177 tagged: 308 with a run with no instrument failure, 70 where every run was a failure of Syntology's instrument)
Page 17 of 22: papers 1,601 to 1,700 of 2,177, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Robust Visual Question Answering: Datasets, Methods, and Future Challenges21 Jul 2023 0 repositories listed
-
A reinforcement learning approach for VQA validation: an application to diabetic macular edema grading19 Jul 2023 0 repositories listed
-
Generative Visual Question Answering18 Jul 2023 0 repositories listed
-
Let's ViCE! Mimicking Human Cognitive Behavior in Image Generation Evaluation18 Jul 2023 0 repositories listed
-
PAT: Parallel Attention Transformer for Visual Question Answering in Vietnamese17 Jul 2023 0 repositories listed
-
A scoping review on multimodal deep learning in biomedical images and texts14 Jul 2023 0 repositories listed
-
Structure Guided Multi-modal Pre-trained Transformer for Knowledge Graph Reasoning6 Jul 2023 0 repositories listed
-
UIT-Saviors at MEDVQA-GI 2023: Improving Multimodal Learning with Image Enhancement for Gastrointestinal Visual Question Answering6 Jul 2023 0 repositories listed
-
Switch-BERT: Learning to Model Multimodal Interactions by Switching Attention and Input25 Jun 2023 0 repositories listed
-
Visual Question Answering in Remote Sensing with Cross-Attention and Multimodal Information Bottleneck25 Jun 2023 0 repositories listed
-
TaCA: Upgrading Your Visual Foundation Model with Task-agnostic Compatible Adapter22 Jun 2023 0 repositories listed
-
AVIS: Autonomous Visual Information Seeking with Large Language Model Agent13 Jun 2023 0 repositories listed
-
Visual Question Answering (VQA) on Images with Superimposed Text13 Jun 2023 0 repositories listed
-
A Survey of Vision-Language Pre-training from the Lens of Multimodal Machine Translation12 Jun 2023 0 repositories listed
-
Knowledge Detection by Relevant Question and Image Attributes in Visual Question Answering8 Jun 2023 0 repositories listed
-
Diversifying Joint Vision-Language Tokenization Learning6 Jun 2023 0 repositories listed
-
Multi-CLIP: Contrastive Vision-Language Pre-training for Question Answering tasks in 3D Scenes4 Jun 2023 0 repositories listed
-
Evaluating the Capabilities of Multi-modal Reasoning Models with Synthetic Task Data1 Jun 2023 0 repositories listed
-
LiT-4-RSVQA: Lightweight Transformer-based Visual Question Answering in Remote Sensing1 Jun 2023 0 repositories listed
-
Overcoming Language Bias in Remote Sensing Visual Question Answering via Adversarial Training1 Jun 2023 0 repositories listed
-
Unveiling Cross Modality Bias in Visual Question Answering: A Causal View with Possible Worlds VQA31 May 2023 0 repositories listed
-
Using Visual Cropping to Enhance Fine-Detail Question Answering of BLIP-Family Models31 May 2023 0 repositories listed
-
Generate then Select: Open-ended Visual Question Answering Guided by World Knowledge30 May 2023 0 repositories listed
-
Mindstorms in Natural Language-Based Societies of Mind26 May 2023 0 repositories listed
-
EmbodiedGPT: Vision-Language Pre-Training via Embodied Chain of Thought24 May 2023 0 repositories listed
-
GRILL: Grounded Vision-language Pre-training via Aligning Text and Image Regions24 May 2023 0 repositories listed
-
Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering24 May 2023 0 repositories listed
-
23 May 2023 0 repositories listed
-
i-Code Studio: A Configurable and Composable Framework for Integrative AI23 May 2023 0 repositories listed
-
Image Manipulation via Multi-Hop Instructions -- A New Dataset and Weakly-Supervised Neuro-Symbolic Approach23 May 2023 0 repositories listed
-
Visual Question Answering: A Survey on Techniques and Common Trends in Recent Literature18 May 2023 0 repositories listed
-
An Empirical Study on the Language Modal in Visual Question Answering17 May 2023 0 repositories listed
-
Probing the Role of Positional Information in Vision-Language Models17 May 2023 0 repositories listed
-
Semantic Composition in Visually Grounded Language Models15 May 2023 0 repositories listed
-
Analysis of Visual Question Answering Algorithms with attention model4 May 2023 0 repositories listed
-
Making the Most of What You Have: Adapting Pre-trained Visual Language Models in the Low-data Regime3 May 2023 0 repositories listed
-
CHIC: Corporate Document for Visual question Answering1 May 2023 0 repositories listed
-
Chain of Thought Prompt Tuning in Vision Language Models16 Apr 2023 0 repositories listed
-
13 Apr 2023 0 repositories listed
-
Advancing Medical Imaging with Language Models: A Journey from N-grams to ChatGPT11 Apr 2023 0 repositories listed
-
Boosting Cross-task Transferability of Adversarial Patches with Visual Relations11 Apr 2023 0 repositories listed
-
CAVL: Learning Contrastive and Adaptive Representations of Vision and Language10 Apr 2023 0 repositories listed
-
Multilingual Augmentation for Robust Visual Question Answering in Remote Sensing Images7 Apr 2023 0 repositories listed
-
Improving Visual Question Answering Models through Robustness Analysis and In-Context Learning with a Chain of Basic Questions6 Apr 2023 0 repositories listed
-
Locate Then Generate: Bridging Vision and Language with Bounding Box for Scene-Text VQA4 Apr 2023 0 repositories listed
-
Q2ATransformer: Improving Medical VQA via an Answer Querying Decoder4 Apr 2023 0 repositories listed
-
SC-ML: Self-supervised Counterfactual Metric Learning for Debiased Visual Question Answering4 Apr 2023 0 repositories listed
-
Instance-Level Trojan Attacks on Visual Question Answering via Adversarial Learning in Neuron Activation Space2 Apr 2023 0 repositories listed
-
Curriculum Learning for Compositional Visual Reasoning27 Mar 2023 0 repositories listed
-
3D Concept Learning and Reasoning from Multi-View Images20 Mar 2023 0 repositories listed
-
FVQA 2.0: Introducing Adversarial Samples into Fact-based Visual Question Answering19 Mar 2023 0 repositories listed
-
Breaking Common Sense: WHOOPS! A Vision-and-Language Benchmark of Synthetic and Compositional Images13 Mar 2023 0 repositories listed
-
Polar-VQA: Visual Question Answering on Remote Sensed Ice sheet Imagery from Polar Region13 Mar 2023 0 repositories listed
-
Vision-Language Models as Success Detectors13 Mar 2023 0 repositories listed
-
Understanding and Constructing Latent Modality Structures in Multi-modal Representation Learning10 Mar 2023 0 repositories listed
-
Toward Unsupervised Realistic Visual Question Answering9 Mar 2023 0 repositories listed
-
Interpretable Visual Question Answering Referring to Outside Knowledge8 Mar 2023 0 repositories listed
-
Graph Neural Networks in Vision-Language Image Understanding: A Survey7 Mar 2023 0 repositories listed
-
Knowledge-Based Counterfactual Queries for Visual Question Answering5 Mar 2023 0 repositories listed
-
VQA with Cascade of Self- and Co-Attention Blocks28 Feb 2023 0 repositories listed
-
Medical visual question answering using joint self-supervised learning25 Feb 2023 0 repositories listed
-
23 Feb 2023 0 repositories listed
-
Reusable Slotwise Mechanisms21 Feb 2023 0 repositories listed
-
Few-shot Multimodal Multitask Multilingual Learning19 Feb 2023 0 repositories listed
-
Interpretable Medical Image Visual Question Answering via Multi-Modal Relationship Graph Learning19 Feb 2023 0 repositories listed
-
Bridge Damage Cause Estimation Using Multiple Images Based on Visual Question Answering18 Feb 2023 0 repositories listed
-
Towards a Unified Model for Generating Answers and Explanations in Visual Question Answering25 Jan 2023 0 repositories listed
-
HRVQA: A Visual Question Answering Benchmark for High-Resolution Aerial Images23 Jan 2023 0 repositories listed
-
Towards Models that Can See and Read18 Jan 2023 0 repositories listed
-
Curriculum Script Distillation for Multilingual Visual Question Answering17 Jan 2023 0 repositories listed
-
Decouple Before Interact: Multi-Modal Prompt Learning for Continual Visual Question Answering1 Jan 2023 0 repositories listed
-
From Images to Textual Prompts: Zero-Shot Visual Question Answering With Frozen Large Language Models1 Jan 2023 0 repositories listed
-
Image as a Foreign Language: BEiT Pretraining for Vision and Vision-Language Tasks1 Jan 2023 0 repositories listed
-
PromptCap: Prompt-Guided Image Captioning for VQA with GPT-31 Jan 2023 0 repositories listed
-
RMLVQA: A Margin Loss Approach for Visual Question Answering With Language Biases1 Jan 2023 0 repositories listed
-
When are Lemons Purple? The Concept Association Bias of Vision-Language Models22 Dec 2022 0 repositories listed
-
UnICLAM:Contrastive Representation Learning with Adversarial Masking for Unified and Interpretable Medical Vision Question Answering21 Dec 2022 0 repositories listed
-
Towards Unsupervised Visual Reasoning: Do Off-The-Shelf Features Know How to Reason?20 Dec 2022 0 repositories listed
-
SceneGATE: Scene-Graph based co-Attention networks for TExt visual question answering16 Dec 2022 0 repositories listed
-
7 Dec 2022 0 repositories listed
-
Compound Tokens: Channel Fusion for Vision-Language Representation Learning2 Dec 2022 0 repositories listed
-
Optimizing Explanations by Network Canonization and Hyperparameter Search30 Nov 2022 0 repositories listed
-
PiggyBack: Pretrained Visual Question Answering Environment for Backing up Non-deep Learning Professionals29 Nov 2022 0 repositories listed
-
Neuro-Symbolic Spatio-Temporal Reasoning28 Nov 2022 0 repositories listed
-
Look, Read and Ask: Learning to Ask Questions by Reading Text in Images23 Nov 2022 0 repositories listed
-
CL-CrossVQA: A Continual Learning Benchmark for Cross-Domain Visual Question Answering19 Nov 2022 0 repositories listed
-
Text-Aware Dual Routing Network for Visual Question Answering17 Nov 2022 0 repositories listed
-
AlignVE: Visual Entailment Recognition Based on Alignment Relations16 Nov 2022 0 repositories listed
-
MF2-MVQA: A Multi-stage Feature Fusion method for Medical Visual Question Answering11 Nov 2022 0 repositories listed
-
9 Nov 2022 0 repositories listed
-
Towards Reasoning-Aware Explainable VQA9 Nov 2022 0 repositories listed
-
Generalization Differences between End-to-End and Neuro-Symbolic Vision-Language Reasoning Systems26 Oct 2022 0 repositories listed
-
Learning by Hallucinating: Vision-Language Pre-training with Weak Supervision24 Oct 2022 0 repositories listed
-
CPL: Counterfactual Prompt Learning for Vision and Language Models19 Oct 2022 0 repositories listed
-
Image Semantic Relation Generation19 Oct 2022 0 repositories listed
-
Aligning MAGMA by Few-Shot Learning and Finetuning18 Oct 2022 0 repositories listed
-
Entity-Focused Dense Passage Retrieval for Outside-Knowledge Visual Question Answering18 Oct 2022 0 repositories listed
-
Multi-Modal Fusion Transformer for Visual Question Answering in Remote Sensing10 Oct 2022 0 repositories listed
-
MAMO: Masked Multimodal Modeling for Fine-Grained Vision-Language Representation Learning9 Oct 2022 0 repositories listed
-
Dual Capsule Attention Mask Network with Mutual Learning for Visual Question Answering1 Oct 2022 0 repositories listed