Browse State-of-the-Art › Video Question Answering › Papers, page 4
Video Question Answering
Papers archive 2025-07-28
archive papers tagged: 460 · with a code link: 250 · where Syntology ran a sample: 124 (107 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (124 of 460 tagged: 107 with a run with no instrument failure, 17 where every run was a failure of Syntology's instrument)
Page 4 of 5: papers 301 to 400 of 460, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Prompting Video-Language Foundation Models with Domain-specific Fine-grained Heuristics for Video Question Answering12 Oct 2024 0 repositories listed
-
Sample then Identify: A General Framework for Risk Control and Assessment in Multimodal Large Language Models10 Oct 2024 0 repositories listed
-
Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization9 Oct 2024 0 repositories listed
-
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition8 Oct 2024 0 repositories listed
-
Frame-Voyager: Learning to Query Frames for Video Large Language Models4 Oct 2024 0 repositories listed
-
3 Oct 2024 0 repositories listed
-
Video DataFlywheel: Resolving the Impossible Data Trinity in Video-Language Understanding29 Sep 2024 0 repositories listed
-
First Place Solution to the Multiple-choice Video QA Track of The Second Perception Test Challenge20 Sep 2024 0 repositories listed
-
Uncertainty-Guided Self-Questioning and Answering for Video-Language Alignment17 Sep 2024 0 repositories listed
-
QTG-VQA: Question-Type-Guided Architectural for VideoQA Systems14 Sep 2024 0 repositories listed
-
Multi-object event graph representation learning for Video Question Answering12 Sep 2024 0 repositories listed
-
Top-down Activity Representation Learning for Video Question Answering12 Sep 2024 0 repositories listed
-
Assessing Modality Bias in Video Question Answering Benchmarks with Multimodal Large Language Models22 Aug 2024 0 repositories listed
-
Continuous Perception Benchmark15 Aug 2024 0 repositories listed
-
LLaVA-Surg: Towards Multimodal Surgical Assistant via Structured Surgical Video Learning15 Aug 2024 0 repositories listed
-
Causal Understanding For Video Question Answering23 Jul 2024 0 repositories listed
-
End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling21 Jul 2024 0 repositories listed
-
VDMA: Video Question Answering with Dynamically Generated Multi-Agents4 Jul 2024 0 repositories listed
-
Align and Aggregate: Compositional Reasoning with Video Alignment and Answer Aggregation for Video Question-Answering3 Jul 2024 0 repositories listed
-
KeyVideoLLM: Towards Large-scale Video Keyframe Selection3 Jul 2024 0 repositories listed
-
The Solution for the ICCV 2023 Perception Test Challenge 2023 -- Task 6 -- Grounded videoQA2 Jul 2024 0 repositories listed
-
Hierarchical Memory for Long Video QA30 Jun 2024 0 repositories listed
-
Zero-Shot Long-Form Video Understanding through Screenplay25 Jun 2024 0 repositories listed
-
Hallucination Mitigation Prompts Long-term Video Understanding17 Jun 2024 0 repositories listed
-
17 Jun 2024 0 repositories listed
-
Video Question Answering for People with Visual Impairments Using an Egocentric 360-Degree Camera30 May 2024 0 repositories listed
-
Backpropagation-Free Multi-modal On-Device Model Adaptation via Cloud-Device Collaboration21 May 2024 0 repositories listed
-
VideoQA-SC: Adaptive Semantic Communication for Video Question Answering17 May 2024 0 repositories listed
-
14 May 2024 0 repositories listed
-
29 Apr 2024 0 repositories listed
-
Pegasus-v1 Technical Report23 Apr 2024 0 repositories listed
-
Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models18 Apr 2024 0 repositories listed
-
9 Apr 2024 0 repositories listed
-
Koala: Key frame-conditioned long video-LLM5 Apr 2024 0 repositories listed
-
Neural-Symbolic VideoQA: Learning Compositional Spatio-Temporal Reasoning for Real-world Video Question Answering5 Apr 2024 0 repositories listed
-
VideoDistill: Language-aware Vision Distillation for Video Question Answering1 Apr 2024 0 repositories listed
-
Ranking Distillation for Open-Ended Video Question Answering with Insufficient Labels21 Mar 2024 0 repositories listed
-
LLMs Meet Long Video: Advancing Long Video Question Answering with An Interactive Visual Adapter in LLMs21 Feb 2024 0 repositories listed
-
Slot-VLM: SlowFast Slots for Video-Language Modeling20 Feb 2024 0 repositories listed
-
VideoPrism: A Foundational Visual Encoder for Video Understanding20 Feb 2024 0 repositories listed
-
BDIQA: A New Dataset for Video Question Answering to Explore Cognitive Reasoning through Theory of Mind12 Feb 2024 0 repositories listed
-
Answering from Sure to Uncertain: Uncertainty-Aware Curriculum Learning for Video Question Answering3 Jan 2024 0 repositories listed
-
Language-aware Visual Semantic Distillation for Video Question Answering1 Jan 2024 0 repositories listed
-
On Scaling Up a Multilingual Vision and Language Model1 Jan 2024 0 repositories listed
-
VISTA-LLAMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens1 Jan 2024 0 repositories listed
-
Cross-Modal Reasoning with Event Correlation for Video Question Answering20 Dec 2023 0 repositories listed
-
Perception Test 2023: A Summary of the First Challenge And Outcome20 Dec 2023 0 repositories listed
-
19 Dec 2023 0 repositories listed
-
12 Dec 2023 0 repositories listed
-
MoVQA: A Benchmark of Versatile Question-Answering for Long-Form Movie Understanding8 Dec 2023 0 repositories listed
-
Retrieval-based Video Language Model for Efficient Long Video Question Answering8 Dec 2023 0 repositories listed
-
VaQuitA: Enhancing Alignment in LLM-Assisted Video Understanding4 Dec 2023 0 repositories listed
-
1 Dec 2023 0 repositories listed
-
E-ViLM: Efficient Video-Language Model via Masked Video Modeling with Semantic Vector-Quantized Tokenizer28 Nov 2023 0 repositories listed
-
Characterizing Video Question Answering with Sparsified Inputs27 Nov 2023 0 repositories listed
-
GPT4Video: A Unified Multimodal Large Language Model for lnstruction-Followed Understanding and Safety-Aware Generation25 Nov 2023 0 repositories listed
-
9 Nov 2023 0 repositories listed
-
Modular Blended Attention Network for Video Question Answering2 Nov 2023 0 repositories listed
-
6 Oct 2023 0 repositories listed
-
5 Sep 2023 0 repositories listed
-
Understanding Video Scenes through Text: Insights from Text-based Video Question Answering4 Sep 2023 0 repositories listed
-
Distraction-free Embeddings for Robust VQA31 Aug 2023 0 repositories listed
-
Redundancy-aware Transformer for Video Question Answering7 Aug 2023 0 repositories listed
-
Keyword-Aware Relative Spatio-Temporal Graph Networks for Video Question Answering25 Jul 2023 0 repositories listed
-
Traffic-Domain Video Question Answering with Automatic Captioning18 Jul 2023 0 repositories listed
-
Read, Look or Listen? What's Needed for Solving a Multimodal Dataset6 Jul 2023 0 repositories listed
-
15 Jun 2023 0 repositories listed
-
Diversifying Joint Vision-Language Tokenization Learning6 Jun 2023 0 repositories listed
-
22 May 2023 0 repositories listed
-
TG-VQA: Ternary Game of Video Question Answering17 May 2023 0 repositories listed
-
Is a Video worth n×n Images? A Highly Efficient Approach to Transformer-based Video Question Answering16 May 2023 0 repositories listed
-
Semantic-aware Dynamic Retrospective-Prospective Reasoning for Event-level Video Question Answering14 May 2023 0 repositories listed
-
VideoOFA: Two-Stage Pre-Training for Video-to-Text Generation4 May 2023 0 repositories listed
-
A Review of Deep Learning for Video Captioning22 Apr 2023 0 repositories listed
-
Language Models are Causal Knowledge Extractors for Zero-shot Video Question Answering7 Apr 2023 0 repositories listed
-
Structured Video-Language Modeling with Temporal Grouping and Spatial Grounding28 Mar 2023 0 repositories listed
-
10 Mar 2023 0 repositories listed
-
Video Question Answering Using CLIP-Guided Visual-Text Attention6 Mar 2023 0 repositories listed
-
STOA-VLP: Spatial-Temporal Modeling of Object and Action for Video-Language Pre-training20 Feb 2023 0 repositories listed
-
27 Jan 2023 0 repositories listed
-
Temporal Perceiving Video-Language Pre-training18 Jan 2023 0 repositories listed
-
Learning Trajectory-Word Alignments for Video-Language Tasks5 Jan 2023 0 repositories listed
-
Discovering the Real Association: Multimodal Causal Reasoning in Video Question Answering1 Jan 2023 0 repositories listed
-
Knowledge Proxy Intervention for Deconfounded Video Question Answering1 Jan 2023 0 repositories listed
-
30 Dec 2022 0 repositories listed
-
9 Dec 2022 0 repositories listed
-
SMAUG: Sparse Masked Autoencoder for Efficient Video-Language Pre-training21 Nov 2022 0 repositories listed
-
Watching the News: Towards VideoQA Models that can Read10 Nov 2022 0 repositories listed
-
LiteVL: Efficient Video-Language Learning with Enhanced Spatial-Temporal Modeling21 Oct 2022 0 repositories listed
-
20 Oct 2022 0 repositories listed
-
Dense but Efficient VideoQA for Intricate Compositional Reasoning19 Oct 2022 0 repositories listed
-
Contrastive Video-Language Learning with Fine-grained Frame Sampling10 Oct 2022 0 repositories listed
-
Locate before Answering: Answer Guided Question Localization for Video Question Answering5 Oct 2022 0 repositories listed
-
In-the-Wild Video Question Answering1 Oct 2022 0 repositories listed
-
15 Sep 2022 0 repositories listed
-
14 Sep 2022 0 repositories listed
-
Frame-Subtitle Self-Supervision for Multi-Modal Video Question Answering8 Sep 2022 0 repositories listed
-
1 Aug 2022 0 repositories listed
-
Dynamic Multistep Reasoning based on Video Scene Graph for Video Question Answering1 Jul 2022 0 repositories listed
-
19 Jun 2022 0 repositories listed