Browse State-of-the-Art › Visual Question Answering › Papers, page 13
Visual Question Answering
Papers archive 2025-07-28
archive papers tagged: 2,177 · with a code link: 1,042 · where Syntology ran a sample: 378 (308 with a run with no instrument failure, 70 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (378 of 2,177 tagged: 308 with a run with no instrument failure, 70 where every run was a failure of Syntology's instrument)
Page 13 of 22: papers 1,201 to 1,300 of 2,177, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Scene Understanding Enabled Semantic Communication with Open Channel Coding24 Jan 2025 0 repositories listed
-
Combining Knowledge Graph and LLMs for Enhanced Zero-shot Visual Question Answering22 Jan 2025 0 repositories listed
-
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!18 Jan 2025 0 repositories listed
-
Double Visual Defense: Adversarial Pre-training and Instruction Tuning for Improving Vision-Language Model Robustness16 Jan 2025 0 repositories listed
-
Dynamic Knowledge Integration for Enhanced Vision-Language Reasoning15 Jan 2025 0 repositories listed
-
Embodied Scene Understanding for Vision Language Models via MetaVQA15 Jan 2025 0 repositories listed
-
SAR Strikes Back: A New Hope for RSVQA14 Jan 2025 0 repositories listed
-
The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering13 Jan 2025 0 repositories listed
-
GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing12 Jan 2025 0 repositories listed
-
Overcoming Language Priors for Visual Question Answering Based on Knowledge Distillation10 Jan 2025 0 repositories listed
-
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding9 Jan 2025 0 repositories listed
-
Feedback-Driven Vision-Language Alignment with Minimal Human Supervision8 Jan 2025 0 repositories listed
-
7 Jan 2025 0 repositories listed
-
Visual question answering: from early developments to recent advances -- a survey7 Jan 2025 0 repositories listed
-
Accounting for Focus Ambiguity in Visual Questions4 Jan 2025 0 repositories listed
-
Interpretable Face Anti-Spoofing: Enhancing Generalization with Multimodal Large Language Models3 Jan 2025 0 repositories listed
-
MoColl: Agent-Based Specific and General Model Collaboration for Image Captioning3 Jan 2025 0 repositories listed
-
CLIP-UP: CLIP-Based Unanswerable Problem Detection for Visual Question Answering2 Jan 2025 0 repositories listed
-
AdaDARE-gamma: Balancing Stability and Plasticity in Multi-modal LLMs through Efficient Adaptation1 Jan 2025 0 repositories listed
-
Alignment, Mining and Fusion: Representation Alignment with Hard Negative Mining and Selective Knowledge Fusion for Medical Visual Question Answering1 Jan 2025 0 repositories listed
-
EfficientLLaVA: Generalizable Auto-Pruning for Large Vision-language Models1 Jan 2025 0 repositories listed
-
JTD-UAV: MLLM-Enhanced Joint Tracking and Description Framework for Anti-UAV Systems1 Jan 2025 0 repositories listed
-
Separation of Powers: On Segregating Knowledge from Observation in LLM-enabled Knowledge-based Visual Question Answering1 Jan 2025 0 repositories listed
-
Probing Visual Language Priors in VLMs31 Dec 2024 0 repositories listed
-
Enhanced Multimodal RAG-LLM for Accurate Visual Question Answering30 Dec 2024 0 repositories listed
-
ErgoChat: a Visual Query System for the Ergonomic Risk Assessment of Construction Workers27 Dec 2024 0 repositories listed
-
Multi-Agents Based on Large Language Models for Knowledge-based Visual Question Answering24 Dec 2024 0 repositories listed
-
TextMatch: Enhancing Image-Text Consistency Through Multimodal Optimization24 Dec 2024 0 repositories listed
-
Cross-Lingual Text-Rich Visual Comprehension: An Information Theory Perspective23 Dec 2024 0 repositories listed
-
FFA Sora, video generation as fundus fluorescein angiography simulator23 Dec 2024 0 repositories listed
-
Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy23 Dec 2024 0 repositories listed
-
Prompting Large Language Models with Rationale Heuristics for Knowledge-based Visual Question Answering22 Dec 2024 0 repositories listed
-
FedPIA -- Permuting and Integrating Adapters leveraging Wasserstein Barycenters for Finetuning Foundation Models in Multi-Modal Federated Learning19 Dec 2024 0 repositories listed
-
A Concept-Centric Approach to Multi-Modality Learning18 Dec 2024 0 repositories listed
-
CPath-Omni: A Unified Multimodal Foundation Model for Patch and Whole Slide Image Analysis in Computational Pathology16 Dec 2024 0 repositories listed
-
Overview of TREC 2024 Medical Video Question Answering (MedVidQA) Track15 Dec 2024 0 repositories listed
-
Damage Assessment after Natural Disasters with UAVs: Semantic Feature Extraction using Deep Learning14 Dec 2024 0 repositories listed
-
Patch-level Sounding Object Tracking for Audio-Visual Question Answering14 Dec 2024 0 repositories listed
-
VLR-Bench: Multilingual Benchmark Dataset for Vision-Language Retrieval Augmented Generation13 Dec 2024 0 repositories listed
-
ViUniT: Visual Unit Tests for More Robust Visual Programming12 Dec 2024 0 repositories listed
-
A Multimodal Social Agent11 Dec 2024 0 repositories listed
-
Barking Up The Syntactic Tree: Enhancing VLM Training with Syntactic Losses11 Dec 2024 0 repositories listed
-
Can We Generate Visual Programs Without Prompting LLMs?11 Dec 2024 0 repositories listed
-
How Vision-Language Tasks Benefit from Large Pre-trained Models: A Survey11 Dec 2024 0 repositories listed
-
9 Dec 2024 0 repositories listed
-
Ranked from Within: Ranking Large Multimodal Models for Visual Question Answering Without Labels9 Dec 2024 0 repositories listed
-
EACO: Enhancing Alignment in Multimodal LLMs via Critical Observation6 Dec 2024 0 repositories listed
-
Findings of the Second BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora6 Dec 2024 0 repositories listed
-
T2I-FactualBench: Benchmarking the Factuality of Text-to-Image Models with Knowledge-Intensive Concepts5 Dec 2024 0 repositories listed
-
CEGI: Measuring the trade-off between efficiency and carbon emissions for SLMs and VLMs3 Dec 2024 0 repositories listed
-
Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey3 Dec 2024 0 repositories listed
-
Understanding the World's Museums through Vision-Language Reasoning2 Dec 2024 0 repositories listed
-
Beyond Logit Lens: Contextual Embeddings for Robust Hallucination Detection & Grounding in VLMs28 Nov 2024 0 repositories listed
-
Sparse Attention Vectors: Generative Multimodal Model Features Are Discriminative Vision-Language Classifiers28 Nov 2024 0 repositories listed
-
Active Data Curation Effectively Distills Large-Scale Multimodal Models27 Nov 2024 0 repositories listed
-
ElectroVizQA: How well do Multi-modal LLMs perform in Electronics Visual Question Answering?27 Nov 2024 0 repositories listed
-
Efficient Multi-modal Large Language Models via Visual Token Grouping26 Nov 2024 0 repositories listed
-
Natural Language Understanding and Inference with MLLM in Visual Question Answering: A Survey26 Nov 2024 0 repositories listed
-
Task Progressive Curriculum Learning for Robust Visual Question Answering26 Nov 2024 0 repositories listed
-
GEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-ray Diagnosis25 Nov 2024 0 repositories listed
-
Text-Guided Coarse-to-Fine Fusion Network for Robust Remote Sensing Visual Question Answering24 Nov 2024 0 repositories listed
-
FINECAPTION: Compositional Image Captioning Focusing on Wherever You Want at Any Granularity23 Nov 2024 0 repositories listed
-
freePruner: A Training-free Approach for Large Multimodal Model Acceleration23 Nov 2024 0 repositories listed
-
ReWind: Understanding Long Videos with Instructed Learnable Memory23 Nov 2024 0 repositories listed
-
21 Nov 2024 0 repositories listed
-
21 Nov 2024 0 repositories listed
-
LaVida Drive: Vision-Text Interaction VLM for Autonomous Driving with Token Selection, Recovery and Enhancement20 Nov 2024 0 repositories listed
-
Uni-Mlip: Unified Self-supervision for Medical Vision Language Pre-training20 Nov 2024 0 repositories listed
-
CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs19 Nov 2024 0 repositories listed
-
Med-2E3: A 2D-Enhanced 3D Medical Multimodal Large Language Model19 Nov 2024 0 repositories listed
-
A Comprehensive Survey on Visual Question Answering Datasets and Algorithms17 Nov 2024 0 repositories listed
-
Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry17 Nov 2024 0 repositories listed
-
Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering17 Nov 2024 0 repositories listed
-
Large Vision-Language Models for Remote Sensing Visual Question Answering16 Nov 2024 0 repositories listed
-
AMXFP4: Taming Activation Outliers with Asymmetric Microscaling Floating-Point for 4-bit LLM Inference15 Nov 2024 0 repositories listed
-
Everything is a Video: Unifying Modalities through Next-Frame Prediction15 Nov 2024 0 repositories listed
-
Visual question answering based evaluation metrics for text-to-image generation15 Nov 2024 0 repositories listed
-
8 Nov 2024 0 repositories listed
-
Integrating Object Detection Modality into Visual Language Model for Enhanced Autonomous Driving Agent8 Nov 2024 0 repositories listed
-
M3DocRAG: Multi-modal Retrieval is What You Need for Multi-page Multi-document Understanding7 Nov 2024 0 repositories listed
-
SaSR-Net: Source-Aware Semantic Representation Network for Enhancing Audio-Visual Question Answering7 Nov 2024 0 repositories listed
-
Seeing is Deceiving: Exploitation of Visual Pathways in Multi-Modal Language Models7 Nov 2024 0 repositories listed
-
NeurIPS 2023 Competition: Privacy Preserving Federated Learning Document VQA6 Nov 2024 0 repositories listed
-
Select2Plan: Training-Free ICL-Based Planning through VQA and Memory Retrieval6 Nov 2024 0 repositories listed
-
From Pixels to Prose: Advancing Multi-Modal Language Models for Remote Sensing5 Nov 2024 0 repositories listed
-
MME-Finance: A Multimodal Finance Benchmark for Expert-level Understanding and Reasoning5 Nov 2024 0 repositories listed
-
Multimodal Commonsense Knowledge Distillation for Visual Question Answering5 Nov 2024 0 repositories listed
-
One VLM to Keep it Learning: Generation and Balancing for Data-free Continual Visual Question Answering4 Nov 2024 0 repositories listed
-
A Visual Question Answering Method for SAR Ship: Breaking the Requirement for Multimodal Dataset Construction and Model Fine-Tuning3 Nov 2024 0 repositories listed
-
Goal-Oriented Semantic Communication for Wireless Visual Question Answering3 Nov 2024 0 repositories listed
-
RS-MoE: Mixture of Experts for Remote Sensing Image Captioning and Visual Question Answering3 Nov 2024 0 repositories listed
-
Designing a Robust Radiology Report Generation System2 Nov 2024 0 repositories listed
-
SimpsonsVQA: Enhancing Inquiry-Based Learning with a Tailored Dataset30 Oct 2024 0 repositories listed
-
GRADE: Quantifying Sample Diversity in Text-to-Image Models29 Oct 2024 0 repositories listed
-
Attention Overlap Is Responsible for The Entity Missing Problem in Text-to-image Diffusion Models!28 Oct 2024 0 repositories listed
-
Efficient Bilinear Attention-based Fusion for Medical Visual Question Answering28 Oct 2024 0 repositories listed
-
Face-MLLM: A Large Face Perception Model28 Oct 2024 0 repositories listed
-
R-LLaVA: Improving Med-VQA Understanding through Visual Region of Interest27 Oct 2024 0 repositories listed
-
GiVE: Guiding Visual Encoder to Perceive Overlooked Information26 Oct 2024 0 repositories listed
-
Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors26 Oct 2024 0 repositories listed