Browse State-of-the-Art › Visual Question Answering (VQA) › Papers, page 13
Visual Question Answering (VQA)
Papers archive 2025-07-28
archive papers tagged: 2,167 · with a code link: 1,039 · where Syntology ran a sample: 359 (287 with a run with no instrument failure, 72 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (359 of 2,167 tagged: 287 with a run with no instrument failure, 72 where every run was a failure of Syntology's instrument)
Page 13 of 22: papers 1,201 to 1,300 of 2,167, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Social-LLaVA: Enhancing Robot Navigation through Human-Language Reasoning in Social Spaces30 Dec 2024 0 repositories listed
-
ESVQA: Perceptual Quality Assessment of Egocentric Spatial Videos29 Dec 2024 0 repositories listed
-
ErgoChat: a Visual Query System for the Ergonomic Risk Assessment of Construction Workers27 Dec 2024 0 repositories listed
-
Not all Views are Created Equal: Analyzing Viewpoint Instabilities in Vision Foundation Models27 Dec 2024 0 repositories listed
-
FineVQ: Fine-Grained User Generated Content Video Quality Assessment26 Dec 2024 0 repositories listed
-
HAUR: Human Annotation Understanding and Recognition Through Text-Heavy Images24 Dec 2024 0 repositories listed
-
Multi-Agents Based on Large Language Models for Knowledge-based Visual Question Answering24 Dec 2024 0 repositories listed
-
TextMatch: Enhancing Image-Text Consistency Through Multimodal Optimization24 Dec 2024 0 repositories listed
-
Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective23 Dec 2024 0 repositories listed
-
Prompting Large Language Models with Rationale Heuristics for Knowledge-based Visual Question Answering22 Dec 2024 0 repositories listed
-
Application of Multimodal Large Language Models in Autonomous Driving21 Dec 2024 0 repositories listed
-
Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage20 Dec 2024 0 repositories listed
-
OnlineVPO: Align Video Diffusion Model with Online Video-Centric Preference Optimization19 Dec 2024 0 repositories listed
-
What makes a good metric? Evaluating automatic metrics for text-to-image consistency18 Dec 2024 0 repositories listed
-
Optimizing Vision-Language Interactions Through Decoder-Only Models14 Dec 2024 0 repositories listed
-
Selective State Space Memory for Large Vision-Language Models13 Dec 2024 0 repositories listed
-
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation13 Dec 2024 0 repositories listed
-
Can We Generate Visual Programs Without Prompting LLMs?11 Dec 2024 0 repositories listed
-
Verb Mirage: Unveiling and Assessing Verb Concept Hallucinations in Multimodal Large Language Models6 Dec 2024 0 repositories listed
-
T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts5 Dec 2024 0 repositories listed
-
AdvDreamer Unveils: Are Vision-Language Models Truly Ready for Real-World 3D Variations?4 Dec 2024 0 repositories listed
-
CEGI: Measuring the trade-off between efficiency and carbon emissions for SLMs and VLMs3 Dec 2024 0 repositories listed
-
WSI-LLaVA: A Multimodal Large Language Model for Whole Slide Image3 Dec 2024 0 repositories listed
-
Perception Test 2024: Challenge Summary and a Novel Hour-Long VideoQA Benchmark29 Nov 2024 0 repositories listed
-
Sparse Attention Vectors: Generative Multimodal Model Features Are Discriminative Vision-Language Classifiers28 Nov 2024 0 repositories listed
-
ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering?27 Nov 2024 0 repositories listed
-
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey26 Nov 2024 0 repositories listed
-
Task Progressive Curriculum Learning for Robust Visual Question Answering26 Nov 2024 0 repositories listed
-
GEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-ray Diagnosis25 Nov 2024 0 repositories listed
-
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models25 Nov 2024 0 repositories listed
-
Enhancing Instruction-Following Capability of Visual-Language Models by Reducing Image Redundancy23 Nov 2024 0 repositories listed
-
ReWind: Understanding Long Videos with Instructed Learnable Memory23 Nov 2024 0 repositories listed
-
Benchmarking Multimodal Models for Ukrainian Language Understanding Across Academic and Cultural Domains22 Nov 2024 0 repositories listed
-
mR²AG: Multimodal Retrieval-Reflection-Augmented Generation for Knowledge-Based VQA22 Nov 2024 0 repositories listed
-
Hints of Prompt: Enhancing Visual Representation for Multimodal LLMs in Autonomous Driving20 Nov 2024 0 repositories listed
-
LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement20 Nov 2024 0 repositories listed
-
Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning Scenarios20 Nov 2024 0 repositories listed
-
Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training20 Nov 2024 0 repositories listed
-
Med-2E3: A 2D-Enhanced 3D Medical Multimodal Large Language Model19 Nov 2024 0 repositories listed
-
A Comprehensive Survey on Visual Question Answering Datasets and Algorithms17 Nov 2024 0 repositories listed
-
F³OCUS -- Federated Finetuning of Vision-Language Foundation Models with Optimal Client Layer Updating Strategy via Multi-objective Meta-Heuristics17 Nov 2024 0 repositories listed
-
Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry17 Nov 2024 0 repositories listed
-
Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering17 Nov 2024 0 repositories listed
-
Visual question answering based evaluation metrics for text-to-image generation15 Nov 2024 0 repositories listed
-
Is Cognition consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding12 Nov 2024 0 repositories listed
-
8 Nov 2024 0 repositories listed
-
NeurIPS 2023 Competition: Privacy Preserving Federated Learning Document VQA6 Nov 2024 0 repositories listed
-
Select2Plan: Training-Free ICL-Based Planning through VQA and Memory Retrieval6 Nov 2024 0 repositories listed
-
MME-Finance: A Multimodal Finance Benchmark for Expert-level Understanding and Reasoning5 Nov 2024 0 repositories listed
-
Multimodal Commonsense Knowledge Distillation for Visual Question Answering5 Nov 2024 0 repositories listed
-
One VLM to Keep it Learning: Generation and Balancing for Data-free Continual Visual Question Answering4 Nov 2024 0 repositories listed
-
A Visual Question Answering Method for SAR Ship: Breaking the Requirement for Multimodal Dataset Construction and Model Fine-Tuning3 Nov 2024 0 repositories listed
-
Goal-Oriented Semantic Communication for Wireless Visual Question Answering3 Nov 2024 0 repositories listed
-
Aggregate-and-Adapt Natural Language Prompts for Downstream Generalization of CLIP31 Oct 2024 0 repositories listed
-
SimpsonsVQA: Enhancing Inquiry-Based Learning with a Tailored Dataset30 Oct 2024 0 repositories listed
-
Attention Overlap Is Responsible for The Entity Missing Problem in Text-to-image Diffusion Models!28 Oct 2024 0 repositories listed
-
Efficient Bilinear Attention-based Fusion for Medical Visual Question Answering28 Oct 2024 0 repositories listed
-
Improving Generalization in Visual Reasoning via Self-Ensemble28 Oct 2024 0 repositories listed
-
R-LLaVA: Improving Med-VQA Understanding through Visual Region of Interest27 Oct 2024 0 repositories listed
-
25 Oct 2024 0 repositories listed
-
Which Client is Reliable?: A Reliable and Personalized Prompt-based Federated Learning for Medical Image Question Answering23 Oct 2024 0 repositories listed
-
Visual Question Answering in Ophthalmology: A Progressive and Practical Perspective22 Oct 2024 0 repositories listed
-
ChitroJera: A Regionally Relevant Visual Question Answering Dataset for Bangla19 Oct 2024 0 repositories listed
-
LLaVA-Ultra: Large Chinese Language and Vision Assistant for Ultrasound19 Oct 2024 0 repositories listed
-
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples18 Oct 2024 0 repositories listed
-
Latent Image and Video Resolution Prediction using Convolutional Neural Networks17 Oct 2024 0 repositories listed
-
RescueADI: Adaptive Disaster Interpretation in Remote Sensing Images with Autonomous Agents17 Oct 2024 0 repositories listed
-
SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image Understanding15 Oct 2024 0 repositories listed
-
Eliminating the Language Bias for Visual Question Answering with fine-grained Causal Intervention14 Oct 2024 0 repositories listed
-
Quality Prediction of AI Generated Images and Videos: Emerging Trends and Opportunities11 Oct 2024 0 repositories listed
-
ViT3D Alignment of LLaMA3: 3D Medical Image Report Generation11 Oct 2024 0 repositories listed
-
Secure Video Quality Assessment Resisting Adversarial Attacks9 Oct 2024 0 repositories listed
-
Beyond Captioning: Task-Specific Prompting for Improved VLM Performance in Mathematical Reasoning8 Oct 2024 0 repositories listed
-
LoGra-Med: Long Context Multi-Graph Alignment for Medical Vision-Language Model3 Oct 2024 0 repositories listed
-
3 Oct 2024 0 repositories listed
-
Backdooring Vision-Language Models with Out-Of-Distribution Data2 Oct 2024 0 repositories listed
-
Why context matters in VQA and Reasoning: Semantic interventions for VLM input modalities2 Oct 2024 0 repositories listed
-
FMBench: Benchmarking Fairness in Multimodal Large Language Models on Medical Tasks1 Oct 2024 0 repositories listed
-
3D-CT-GPT: Generating 3D Radiology Reports through Integration of Large Vision-Language Models28 Sep 2024 0 repositories listed
-
TrojVLM: Backdoor Attack Against Vision Language Models28 Sep 2024 0 repositories listed
-
Visual Question Decomposition on Multimodal Large Language Models28 Sep 2024 0 repositories listed
-
Charting the Future: Using Chart Question-Answering for Scalable Evaluation of LLM-Driven Data Visualizations27 Sep 2024 0 repositories listed
-
DARE: Diverse Visual Question Answering with Robustness Evaluation26 Sep 2024 0 repositories listed
-
ZALM3: Zero-Shot Enhancement of Vision-Language Alignment via In-Context Information in Multi-Turn Multimodal Medical Dialogue26 Sep 2024 0 repositories listed
-
Advancing Video Quality Assessment for AIGC23 Sep 2024 0 repositories listed
-
Detect, Describe, Discriminate: Moving Beyond VQA for MLLM Evaluation23 Sep 2024 0 repositories listed
-
@Bench: Benchmarking Vision-Language Models for Human-centered Assistive Technology21 Sep 2024 0 repositories listed
-
Sparks of Artificial General Intelligence(AGI) in Semiconductor Material Science: Early Explorations into the Next Frontier of Generative AI-Assisted Electron Micrograph Analysis17 Sep 2024 0 repositories listed
-
QTG-VQA: Question-Type-Guided Architectural for VideoQA Systems14 Sep 2024 0 repositories listed
-
Learning to Compress Contexts for Efficient Knowledge-based Visual Question Answering11 Sep 2024 0 repositories listed
-
Look, Learn and Leverage (L³): Mitigating Visual-Domain Shift and Discovering Intrinsic Relations via Symbolic Alignment30 Aug 2024 0 repositories listed
-
Retrieval-Augmented Natural Language Reasoning for Explainable Visual Question Answering30 Aug 2024 0 repositories listed
-
M4CXR: Exploring Multi-task Potentials of Multi-modal Large Language Models for Chest X-ray Interpretation29 Aug 2024 0 repositories listed
-
Can SAR improve RSVQA performance?28 Aug 2024 0 repositories listed
-
Can Visual Language Models Replace OCR-Based Visual Question Answering Pipelines in Production? A Case Study in Retail28 Aug 2024 0 repositories listed
-
Multi-Modal Instruction-Tuning Small-Scale Language-and-Vision Assistant for Semiconductor Electron Micrograph Analysis27 Aug 2024 0 repositories listed
-
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis27 Aug 2024 0 repositories listed
-
LMM-VQA: Advancing Video Quality Assessment with Large Multimodal Models26 Aug 2024 0 repositories listed
-
Towards Human-Level Understanding of Complex Process Engineering Schematics: A Pedagogical, Introspective Multi-Agent Framework for Open-Domain Question Answering24 Aug 2024 0 repositories listed
-
Foundational Model for Electron Micrograph Analysis: Instruction-Tuning Small-Scale Language-and-Vision Assistant for Enterprise Adoption23 Aug 2024 0 repositories listed