Browse State-of-the-Art › Visual Question Answering (VQA) › Papers, page 14
Visual Question Answering (VQA)
Papers archive 2025-07-28
archive papers tagged: 2,167 · with a code link: 1,039 · where Syntology ran a sample: 359 (287 with a run with no instrument failure, 72 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (359 of 2,167 tagged: 287 with a run with no instrument failure, 72 where every run was a failure of Syntology's instrument)
Page 14 of 22: papers 1,301 to 1,400 of 2,167, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Vintern-1B: An Efficient Multimodal Large Language Model for Vietnamese22 Aug 2024 0 repositories listed
-
Swarm Intelligence in Geo-Localization: A Multi-Agent Large Vision-Language Model Collaborative Framework21 Aug 2024 0 repositories listed
-
Beyond the Hype: A dispassionate look at vision-language models in medical scenario16 Aug 2024 0 repositories listed
-
Enhancing Visual Question Answering through Ranking-Based Hybrid Training and Multimodal Fusion14 Aug 2024 0 repositories listed
-
Subjective and Objective Quality Assessment of Rendered Human Avatar Videos in Virtual Reality13 Aug 2024 0 repositories listed
-
Long-Form Answers to Visual Questions from Blind and Low Vision People12 Aug 2024 0 repositories listed
-
Efficient Quantum Gradient and Higher-order Derivative Estimation via Generalized Hadamard Test10 Aug 2024 0 repositories listed
-
Revisiting Multi-Modal LLM Evaluation9 Aug 2024 0 repositories listed
-
Prompting Medical Large Vision-Language Models to Diagnose Pathologies by Visual Question Answering31 Jul 2024 0 repositories listed
-
Benchmarking Multi-dimensional AIGC Video Quality Assessment: A Dataset and Unified Model31 Jul 2024 0 repositories listed
-
SimpleLLM4AD: An End-to-End Vision-Language Model with Graph Visual Question Answering for Autonomous Driving31 Jul 2024 0 repositories listed
-
Highly Efficient No-reference 4K Video Quality Assessment with Full-Pixel Covering Sampling and Training Strategy30 Jul 2024 0 repositories listed
-
Pyramid Coder: Hierarchical Code Generator for Compositional Visual Question Answering30 Jul 2024 0 repositories listed
-
Take A Step Back: Rethinking the Two Stages in Visual Reasoning29 Jul 2024 0 repositories listed
-
Improved Few-Shot Image Classification Through Multiple-Choice Questions23 Jul 2024 0 repositories listed
-
Exploring the Effectiveness of Object-Centric Representations in Visual Question Answering: Comparative Insights with Foundation Models22 Jul 2024 0 repositories listed
-
EchoSight: Advancing Visual-Language Models with Wiki Knowledge17 Jul 2024 0 repositories listed
-
Multimodal Reranking for Knowledge-Intensive Visual Question Answering17 Jul 2024 0 repositories listed
-
TM-PATHVQA:90000+ Textless Multilingual Questions for Medical Visual Question Answering16 Jul 2024 0 repositories listed
-
Extracting Training Data from Document-Based VQA Models11 Jul 2024 0 repositories listed
-
Segmentation-guided Attention for Visual Question Answering from Remote Sensing Images11 Jul 2024 0 repositories listed
-
VQA-Diff: Exploiting VQA and Diffusion for Zero-Shot Image-to-3D Vehicle Asset Generation in Autonomous Driving9 Jul 2024 0 repositories listed
-
BACON: Improving Clarity of Image Captions via Bag-of-Concept Graphs3 Jul 2024 0 repositories listed
-
MindBench: A Comprehensive Benchmark for Mind Map Structure Recognition and Analysis3 Jul 2024 0 repositories listed
-
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness2 Jul 2024 0 repositories listed
-
D-Rax: Domain-specific Radiologic assistant leveraging multi-modal data and eXpert model predictions2 Jul 2024 0 repositories listed
-
https://arxiv.org/abs/2407.006342 Jul 2024 0 repositories listed
-
Hierarchical Memory for Long Video QA30 Jun 2024 0 repositories listed
-
SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs28 Jun 2024 0 repositories listed
-
Disentangling Knowledge-based and Visual Reasoning by Question Decomposition in KB-VQA27 Jun 2024 0 repositories listed
-
RAVEN: Multitask Retrieval Augmented Vision-Language Learning27 Jun 2024 0 repositories listed
-
On the Role of Visual Grounding in VQA26 Jun 2024 0 repositories listed
-
Losing Visual Needles in Image Haystacks: Vision Language Models are Easily Distracted in Short and Long Contexts24 Jun 2024 0 repositories listed
-
MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs24 Jun 2024 0 repositories listed
-
Priorformer: A UGC-VQA Method with content and distortion priors24 Jun 2024 0 repositories listed
-
Tri-VQA: Triangular Reasoning Medical Visual Question Answering for Multi-Attribute Analysis21 Jun 2024 0 repositories listed
-
Precision Empowers, Excess Distracts: Visual Question Answering With Dynamically Infused Knowledge In Language Models14 Jun 2024 0 repositories listed
-
Optimizing Visual Question Answering Models for Driving: Bridging the Gap Between Human and Machine Attention Patterns13 Jun 2024 0 repositories listed
-
CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark10 Jun 2024 0 repositories listed
-
Understanding Information Storage and Transfer in Multi-modal Large Language Models6 Jun 2024 0 repositories listed
-
Translation Deserves Better: Analyzing Translation Artifacts in Cross-lingual Visual Question Answering4 Jun 2024 0 repositories listed
-
Mixture of Rationale: Multi-Modal Reasoning Mixture for Visual Question Answering3 Jun 2024 0 repositories listed
-
Selectively Answering Visual Questions3 Jun 2024 0 repositories listed
-
VQA Training Sets are Self-play Environments for Generating Few-shot Pools30 May 2024 0 repositories listed
-
Evaluating Zero-Shot GPT-4V Performance on 3D Visual Question Answering Benchmarks29 May 2024 0 repositories listed
-
PTM-VQA: Efficient Video Quality Assessment Leveraging Diverse PreTrained Models from the Wild28 May 2024 0 repositories listed
-
Privacy-Aware Visual Language Models27 May 2024 0 repositories listed
-
SearchLVLMs: A Plug-and-Play Framework for Augmenting Large Vision-Language Models by Searching Up-to-Date Internet Knowledge23 May 2024 0 repositories listed
-
Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions18 May 2024 0 repositories listed
-
EyeFound: A Multimodal Generalist Foundation Model for Ophthalmic Imaging18 May 2024 0 repositories listed
-
StackOverflowVQA: Stack Overflow Visual Question Answering Dataset17 May 2024 0 repositories listed
-
RMT-BVQA: Recurrent Memory Transformer-based Blind Video Quality Assessment for Enhanced Video Content14 May 2024 0 repositories listed
-
Realizing Visual Question Answering for Education: GPT-4V as a Multimodal AI12 May 2024 0 repositories listed
-
Is the House Ready For Sleeptime? Generating and Evaluating Situational Queries for Embodied Question Answering8 May 2024 0 repositories listed
-
Advancing Multimodal Medical Capabilities of Gemini6 May 2024 0 repositories listed
-
VSA4VQA: Scaling a Vector Symbolic Architecture to Visual Question Answering on Natural Images6 May 2024 0 repositories listed
-
Beyond Human Vision: The Role of Large Vision Language Models in Microscope Image Analysis1 May 2024 0 repositories listed
-
Enhanced Textual Feature Extraction for Visual Question Answering: A Simple Convolutional Approach1 May 2024 0 repositories listed
-
Visual Fact Checker: Enabling High-Fidelity Detailed Caption Generation30 Apr 2024 0 repositories listed
-
NTIRE 2024 Quality Assessment of AI-Generated Content Challenge25 Apr 2024 0 repositories listed
-
Fusion of Domain-Adapted Vision and Language Models for Medical Visual Question Answering24 Apr 2024 0 repositories listed
-
Exploring Diverse Methods in Visual Question Answering21 Apr 2024 0 repositories listed
-
Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning19 Apr 2024 0 repositories listed
-
PDF-MVQA: A Dataset for Multimodal Information Retrieval in PDF-based Visual Question Answering19 Apr 2024 0 repositories listed
-
TextSquare: Scaling up Text-Centric Visual Instruction Tuning19 Apr 2024 0 repositories listed
-
Unified Scene Representation and Reconstruction for 3D Large Language Models19 Apr 2024 0 repositories listed
-
MedThink: Explaining Medical Visual Question Answering via Multimodal Decision-Making Rationale18 Apr 2024 0 repositories listed
-
Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models18 Apr 2024 0 repositories listed
-
Find The Gap: Knowledge Base Reasoning For Visual Question Answering16 Apr 2024 0 repositories listed
-
BRAVE: Broadening the visual encoding of vision-language models10 Apr 2024 0 repositories listed
-
9 Apr 2024 0 repositories listed
-
HAMMR: HierArchical MultiModal React agents for generic VQA8 Apr 2024 0 repositories listed
-
Study of the effect of Sharpness on Blind Video Quality Assessment6 Apr 2024 0 repositories listed
-
BuDDIE: A Business Document Dataset for Multi-task Information Extraction5 Apr 2024 0 repositories listed
-
TinyVQA: Compact Multimodal Deep Neural Network for Visual Question Answering on Resource-Constrained Devices4 Apr 2024 0 repositories listed
-
Detect2Interact: Localizing Object Key Field in Visual Question Answering (VQA) with LLMs1 Apr 2024 0 repositories listed
-
A Gaze-grounded Visual Question Answering Dataset for Clarifying Ambiguous Japanese Questions26 Mar 2024 0 repositories listed
-
Visual Hallucination: Definition, Quantification, and Prescriptive Remediations26 Mar 2024 0 repositories listed
-
Synthesize Step-by-Step: Tools, Templates and LLMs as Data Generators for Reasoning-Based Chart VQA25 Mar 2024 0 repositories listed
-
Surgical-LVLM: Learning to Adapt Large Vision-Language Model for Grounded Visual Question Answering in Robotic Surgery22 Mar 2024 0 repositories listed
-
AGFSync: Leveraging AI-Generated Feedback for Preference Optimization in Text-to-Image Generation20 Mar 2024 0 repositories listed
-
Multi-Modal Hallucination Control by Visual Information Grounding20 Mar 2024 0 repositories listed
-
WoLF: Wide-scope Large Language Model Framework for CXR Understanding19 Mar 2024 0 repositories listed
-
FlexCap: Describe Anything in Images in Controllable Detail18 Mar 2024 0 repositories listed
-
SpatialPIN: Enhancing Spatial Reasoning Capabilities of Vision-Language Models through Prompting and Interacting 3D Priors18 Mar 2024 0 repositories listed
-
Few-Shot VQA with Frozen LLMs: A Tale of Two Approaches17 Mar 2024 0 repositories listed
-
Few-Shot Image Classification and Segmentation as Visual Question Answering Using Vision-Language Models15 Mar 2024 0 repositories listed
-
Knowledge Condensation and Reasoning for Knowledge-based VQA15 Mar 2024 0 repositories listed
-
Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning15 Mar 2024 0 repositories listed
-
UniCode: Learning a Unified Codebook for Multimodal Large Language Models14 Mar 2024 0 repositories listed
-
8 Mar 2024 0 repositories listed
-
SnapNTell: Enhancing Entity-Centric Visual Question Answering with Retrieval Augmented Multimodal LLM7 Mar 2024 0 repositories listed
-
CLEVR-POC: Reasoning-Intensive Visual Question Answering in Partially Observable Environments5 Mar 2024 0 repositories listed
-
Enhancing Generalization in Medical Visual Question Answering Tasks via Gradient-Guided Model Perturbation5 Mar 2024 0 repositories listed
-
ArcSin: Adaptive ranged cosine Similarity injected noise for Language-Driven Visual Tasks27 Feb 2024 0 repositories listed
-
A Spectrum Evaluation Benchmark for Medical Multi-Modal Large Language Models17 Feb 2024 0 repositories listed
-
PaLM2-VAdapter: Progressively Aligned Language Model Makes a Strong Vision-language Adapter16 Feb 2024 0 repositories listed
-
16 Feb 2024 0 repositories listed Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
15 Feb 2024 0 repositories listed
-
Prompt-based Personalized Federated Learning for Medical Visual Question Answering15 Feb 2024 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.