Browse State-of-the-Art › Visual Question Answering › Papers, page 15
Visual Question Answering
Papers archive 2025-07-28
archive papers tagged: 2,177 · with a code link: 1,042 · where Syntology ran a sample: 378 (308 with a run with no instrument failure, 70 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (378 of 2,177 tagged: 308 with a run with no instrument failure, 70 where every run was a failure of Syntology's instrument)
Page 15 of 22: papers 1,401 to 1,500 of 2,177, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
24 Jun 2024 0 repositories listed
-
GPT-4V Explorations: Mining Autonomous Driving24 Jun 2024 0 repositories listed
-
MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs24 Jun 2024 0 repositories listed
-
MR-MLLM: Mutual Reinforcement of Multimodal Comprehension and Vision Perception22 Jun 2024 0 repositories listed
-
Tri-VQA: Triangular Reasoning Medical Visual Question Answering for Multi-Attribute Analysis21 Jun 2024 0 repositories listed
-
Does Object Grounding Really Reduce Hallucination of Large Vision-Language Models?20 Jun 2024 0 repositories listed
-
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning17 Jun 2024 0 repositories listed
-
Program Synthesis Benchmark for Visual Programming in XLogoOnline Environment17 Jun 2024 0 repositories listed
-
Detecting and Evaluating Medical Hallucinations in Large Vision Language Models14 Jun 2024 0 repositories listed
-
Precision Empowers, Excess Distracts: Visual Question Answering With Dynamically Infused Knowledge In Language Models14 Jun 2024 0 repositories listed
-
SHMamba: Structured Hyperbolic State Space Model for Audio-Visual Question Answering14 Jun 2024 0 repositories listed
-
Optimizing Visual Question Answering Models for Driving: Bridging the Gap Between Human and Machine Attention Patterns13 Jun 2024 0 repositories listed
-
DistilDoc: Knowledge Distillation for Visually-Rich Document Applications12 Jun 2024 0 repositories listed
-
12 Jun 2024 0 repositories listed
-
CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark10 Jun 2024 0 repositories listed
-
Solution for SMART-101 Challenge of CVPR Multi-modal Algorithmic Reasoning Task 202410 Jun 2024 0 repositories listed
-
7 Jun 2024 0 repositories listed
-
6 Jun 2024 0 repositories listed
-
6 Jun 2024 0 repositories listed
-
Understanding Information Storage and Transfer in Multi-modal Large Language Models6 Jun 2024 0 repositories listed
-
Balancing Performance and Efficiency in Zero-shot Robotic Navigation5 Jun 2024 0 repositories listed
-
Story Generation from Visual Inputs: Techniques, Related Tasks, and Challenges4 Jun 2024 0 repositories listed
-
Translation Deserves Better: Analyzing Translation Artifacts in Cross-lingual Visual Question Answering4 Jun 2024 0 repositories listed
-
Mixture of Rationale: Multi-Modal Reasoning Mixture for Visual Question Answering3 Jun 2024 0 repositories listed
-
Selectively Answering Visual Questions3 Jun 2024 0 repositories listed
-
Uncovering Bias in Large Vision-Language Models at Scale with Counterfactuals30 May 2024 0 repositories listed
-
Video Question Answering for People with Visual Impairments Using an Egocentric 360-Degree Camera30 May 2024 0 repositories listed
-
VQA Training Sets are Self-play Environments for Generating Few-shot Pools30 May 2024 0 repositories listed
-
Evaluating Zero-Shot GPT-4V Performance on 3D Visual Question Answering Benchmarks29 May 2024 0 repositories listed
-
MetaToken: Detecting Hallucination in Image Descriptions by Meta Classification29 May 2024 0 repositories listed
-
28 May 2024 0 repositories listed Syntology 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
28 May 2024 0 repositories listed
-
Prompt-Aware Adapter: Towards Learning Adaptive Visual Tokens for Multimodal Large Language Models24 May 2024 0 repositories listed
-
23 May 2024 0 repositories listed
-
SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge23 May 2024 0 repositories listed
-
22 May 2024 0 repositories listed
-
Inquire, Interact, and Integrate: A Proactive Agent Collaborative Framework for Zero-Shot Multimodal Medical Reasoning19 May 2024 0 repositories listed
-
EyeFound: A Multimodal Generalist Foundation Model for Ophthalmic Imaging18 May 2024 0 repositories listed
-
StackOverflowVQA: Stack Overflow Visual Question Answering Dataset17 May 2024 0 repositories listed
-
CLIP-Powered TASS: Target-Aware Single-Stream Network for Audio-Visual Question Answering13 May 2024 0 repositories listed
-
Realizing Visual Question Answering for Education: GPT-4V as a Multimodal AI12 May 2024 0 repositories listed
-
Is the House Ready For Sleeptime? Generating and Evaluating Situational Queries for Embodied Question Answering8 May 2024 0 repositories listed
-
Advancing Multimodal Medical Capabilities of Gemini6 May 2024 0 repositories listed
-
Language-Image Models with 3D Understanding6 May 2024 0 repositories listed
-
VSA4VQA: Scaling a Vector Symbolic Architecture to Visual Question Answering on Natural Images6 May 2024 0 repositories listed
-
Beyond Human Vision: The Role of Large Vision Language Models in Microscope Image Analysis1 May 2024 0 repositories listed
-
CREPE: Coordinate-Aware End-to-End Document Parser1 May 2024 0 repositories listed
-
Enhanced Textual Feature Extraction for Visual Question Answering: A Simple Convolutional Approach1 May 2024 0 repositories listed
-
Efficiency in Focus: LayerNorm as a Catalyst for Fine-tuning Medical Visual Language Pre-trained Models25 Apr 2024 0 repositories listed
-
Fusion of Domain-Adapted Vision and Language Models for Medical Visual Question Answering24 Apr 2024 0 repositories listed
-
Grounded Knowledge-Enhanced Medical VLP for Chest X-Ray23 Apr 2024 0 repositories listed
-
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs23 Apr 2024 0 repositories listed
-
WangLab at MEDIQA-M3G 2024: Multimodal Medical Answer Generation using Large Language Models22 Apr 2024 0 repositories listed
-
Exploring Diverse Methods in Visual Question Answering21 Apr 2024 0 repositories listed
-
Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning19 Apr 2024 0 repositories listed
-
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering19 Apr 2024 0 repositories listed
-
TextSquare: Scaling up Text-Centric Visual Instruction Tuning19 Apr 2024 0 repositories listed
-
MedThink: Explaining Medical Visual Question Answering via Multimodal Decision-Making Rationale18 Apr 2024 0 repositories listed
-
Consistency and Uncertainty: Identifying Unreliable Responses From Black-Box Vision-Language Models for Selective Visual Question Answering16 Apr 2024 0 repositories listed
-
Find The Gap: Knowledge Base Reasoning For Visual Question Answering16 Apr 2024 0 repositories listed
-
HOI-Ref: Hand-Object Interaction Referral in Egocentric Vision15 Apr 2024 0 repositories listed
-
9 Apr 2024 0 repositories listed
-
HAMMR: HierArchical MultiModal React agents for generic VQA8 Apr 2024 0 repositories listed
-
Self-Training Large Language Models for Improved Visual Program Synthesis With Visual Reinforcement6 Apr 2024 0 repositories listed
-
BuDDIE: A Business Document Dataset for Multi-task Information Extraction5 Apr 2024 0 repositories listed
-
TinyVQA: Compact Multimodal Deep Neural Network for Visual Question Answering on Resource-Constrained Devices4 Apr 2024 0 repositories listed
-
Enhancing Human-Computer Interaction in Chest X-ray Analysis using Vision and Language Model with Eye Gaze Patterns3 Apr 2024 0 repositories listed
-
Detect2Interact: Localizing Object Key Field in Visual Question Answering (VQA) with LLMs1 Apr 2024 0 repositories listed
-
Uncovering Bias in Large Vision-Language Models with Counterfactuals29 Mar 2024 0 repositories listed
-
A Gaze-grounded Visual Question Answering Dataset for Clarifying Ambiguous Japanese Questions26 Mar 2024 0 repositories listed
-
Visual Hallucination: Definition, Quantification, and Prescriptive Remediations26 Mar 2024 0 repositories listed
-
PropTest: Automatic Property Testing for Improved Visual Programming25 Mar 2024 0 repositories listed
-
Synthesize Step-by-Step: Tools, Templates and LLMs as Data Generators for Reasoning-Based Chart VQA25 Mar 2024 0 repositories listed
-
Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery22 Mar 2024 0 repositories listed
-
MyVLM: Personalizing VLMs for User-Specific Queries21 Mar 2024 0 repositories listed
-
Improved Baselines for Data-efficient Perceptual Augmentation of LLMs20 Mar 2024 0 repositories listed
-
20 Mar 2024 0 repositories listed
-
As Firm As Their Foundations: Can open-sourced foundation models be used to create adversarial examples for downstream tasks?19 Mar 2024 0 repositories listed
-
WoLF: Wide-scope Large Language Model Framework for CXR Understanding19 Mar 2024 0 repositories listed
-
Can LLMs Generate Human-Like Wayfinding Instructions? Towards Platform-Agnostic Embodied Instruction Synthesis18 Mar 2024 0 repositories listed
-
FlexCap: Describe Anything in Images in Controllable Detail18 Mar 2024 0 repositories listed
-
SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors18 Mar 2024 0 repositories listed
-
Few-Shot VQA with Frozen LLMs: A Tale of Two Approaches17 Mar 2024 0 repositories listed
-
Few-Shot Image Classification and Segmentation as Visual Question Answering Using Vision-Language Models15 Mar 2024 0 repositories listed
-
Knowledge Condensation and Reasoning for Knowledge-based VQA15 Mar 2024 0 repositories listed
-
Parameter Efficient Reinforcement Learning from Human Feedback15 Mar 2024 0 repositories listed
-
14 Mar 2024 0 repositories listed
-
VisionGPT: Vision-Language Understanding Agent Using Generalized Multimodal Framework14 Mar 2024 0 repositories listed
-
13 Mar 2024 0 repositories listed
-
Fine-tuning Large Language Models with Sequential Instructions12 Mar 2024 0 repositories listed
-
Mitigating the Impact of Attribute Editing on Face Recognition12 Mar 2024 0 repositories listed
-
SnapNTell: Enhancing Entity-Centric Visual Question Answering with Retrieval Augmented Multimodal LLM7 Mar 2024 0 repositories listed
-
CLEVR-POC: Reasoning-Intensive Visual Question Answering in Partially Observable Environments5 Mar 2024 0 repositories listed
-
Enhancing Generalization in Medical Visual Question Answering Tasks via Gradient-Guided Model Perturbation5 Mar 2024 0 repositories listed
-
Modeling Collaborator: Enabling Subjective Vision Classification With Minimal Human Effort via LLM Tool-Use5 Mar 2024 0 repositories listed
-
MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting5 Mar 2024 0 repositories listed
-
3 Mar 2024 0 repositories listed
-
A Cognitive Evaluation Benchmark of Image Reasoning and Description for Large Vision-Language Models28 Feb 2024 0 repositories listed
-
ArcSin: Adaptive ranged cosine Similarity injected noise for Language-Driven Visual Tasks27 Feb 2024 0 repositories listed
-
VCD: Knowledge Base Guided Visual Commonsense Discovery in Images27 Feb 2024 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.