Browse State-of-the-Art › Video Question Answering
Video Question Answering
250 papers with code · 28 benchmarks · 43 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
28 leaderboard tables shown for this task, 28 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 28 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
43 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 43 until expanded.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 250 papers with code (460 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
9 Feb 2021 16 repositories listed Syntology ran 35 of 43 samples · 8 unverified · 14 pointer-only (licence)We present a convolution-free approach to video classification built exclusively on self-attention over space and time.
-
17 Apr 2023 13 repositories listed Syntology ran 16 of 51 samples · 35 unverifiedInstruction tuning large language models (LLMs) using machine-generated instruction-following data has improved zero-shot capabilities on new tasks, but the idea is less explored in the multimodal field.
-
18 Sep 2024 8 repositories listed Syntology ran 8 of 12 samples · 4 unverifiedWe present the Qwen2-VL Series, an advanced upgrade of the previous Qwen-VL models that redefines the conventional predetermined-resolution approach in visual processing.
-
16 Nov 2023 6 repositories listed Syntology ran 4 of 7 samples · 3 unverified · 1 pointer-only (licence)In this work, we unify visual representation into the language feature space to advance the foundational LLM towards a unified LVLM.
-
20 Apr 2023 6 repositories listedOur work, for the first time, uncovers that properly aligning the visual features with an advanced large language model can possess numerous advanced multi-modal abilities demonstrated by GPT-4, such as detailed image…
-
29 Apr 2022 5 repositories listed Syntology ran 18 of 24 samples · 6 unverified · 7 pointer-only (licence)Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research.
-
17 Jun 2024 4 repositories listed Syntology ran 20 of 30 samples · 10 unverified · 15 pointer-only (licence)Iterative self-improvement, a concept extending beyond personal growth, has found powerful applications in machine learning, particularly in transforming weak models into strong ones.
-
14 Nov 2023 4 repositories listedLarge language models have demonstrated impressive universal capabilities across a wide range of open-ended tasks and have extended their utility to encompass multimodal conversations.
-
5 Jun 2023 4 repositories listed Syntology ran 18 of 25 samples · 7 unverified · 9 pointer-only (licence)We present Video-LLaMA a multi-modal framework that empowers Large Language Models (LLMs) with the capability of understanding both visual and auditory content in the video.
-
11 May 2023 4 repositories listedLarge-scale pre-training and instruction tuning have been successful at creating general-purpose language models with broad competence.
-
Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning25 Mar 2023 4 repositories listed Syntology ran 11 of 16 samples · 5 unverifiedContrastive learning-based video-language representation learning approaches, e.
-
1 Feb 2023 4 repositories listed Syntology ran 9 of 19 samples · 10 unverifiedIn contrast to predominant paradigms of solely relying on sequence-to-sequence generation or encoder-based instance discrimination, mPLUG-2 introduces a multi-module composition network by sharing common universal…
-
21 Nov 2022 4 repositories listed Syntology ran 3 of 4 samples · 1 unverifiedMost video-and-language representation learning approaches employ contrastive learning, e.
-
5 Sep 2018 4 repositories listed Syntology ran 3 of 8 samples · 5 unverifiedRecent years have witnessed an increasing interest in image-based question-answering (QA) tasks.
-
10 Jul 2024 3 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 4 pointer-only (licence)To this end, we introduce LLaVA-NeXT-Interleave, which simultaneously tackles Multi-image, Multi-frame (video), Multi-view (3D), and Multi-patch (single-image) scenarios in LMMs.
-
11 Jun 2024 3 repositories listed Syntology ran 7 of 17 samples · 10 unverifiedIn this paper, we present the VideoLLaMA 2, a set of Video Large Language Models (Video-LLMs) designed to enhance spatial-temporal modeling and audio understanding in video and audio-oriented tasks.
-
28 Nov 2023 3 repositories listed Syntology ran 7 of 10 samples · 3 unverifiedWith the rapid development of Multi-modal Large Language Models (MLLMs), a number of diagnostic benchmarks have recently emerged to evaluate the comprehension capabilities of these models.
-
28 Apr 2023 3 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)This strategy effectively alleviates the interference between the two tasks of image-text alignment and instruction following and achieves strong multi-modal reasoning with only a small-scale image-text and instruction…
-
16 Jun 2022 3 repositories listed Syntology ran 14 of 34 samples · 20 unverified · 1 pointer-only (licence)Manual annotation of question and answers for videos, however, is tedious and prohibits scalability.
-
29 Mar 2021 3 repositories listed Syntology ran 1 of 9 samples · 8 unverified · 2 pointer-only (licence)In this paper, we create a novel dataset, SUTD-TrafficQA (Traffic Question Answering), which takes the form of video QA based on the collected 10, 080 in-the-wild videos and annotated 62, 535 QA pairs, for benchmarking…
-
1 May 2020 3 repositories listed Syntology ran 5 of 13 samples · 8 unverified · 8 pointer-only (licence)We present HERO, a novel framework for large-scale video+language omni-representation learning.
-
25 Apr 2019 3 repositories listed Syntology ran 6 of 13 samples · 7 unverifiedWe present the task of Spatio-Temporal Video Question Answering, which requires intelligent systems to simultaneously retrieve relevant moments and detect referenced visual concepts (people and objects) to answer…
-
8 May 2015 3 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 2 pointer-only (licence)A suite of baseline results on this new dataset are also presented.
-
5 Dec 2024 2 repositories listedThis paper introduces NVILA, a family of open VLMs designed to optimize both efficiency and accuracy.
-
6 Aug 2024 2 repositories listedWe present LLaVA-OneVision, a family of open large multimodal models (LMMs) developed by consolidating our insights into data, models, and visual representations in the LLaVA-NeXT blog series.
-
24 Jun 2024 2 repositories listed Syntology ran 5 of 5 samples · 0 unverified · 5 pointer-only (licence)By simply extrapolating the context length of the language backbone, we enable LMMs to comprehend orders of magnitude more visual tokens without any video training.
-
4 Apr 2024 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedThis paper introduces MiniGPT4-Video, a multimodal Large Language Model (LLM) designed specifically for video understanding.
-
22 Mar 2024 2 repositories listedWe introduce InternVideo2, a new family of video foundation models (ViFM) that achieve the state-of-the-art results in video recognition, video-text tasks, and video-centric dialogue.
-
25 Feb 2024 2 repositories listed Syntology ran 5 of 8 samples · 3 unverifiedOur framework significantly enhances the temporal capabilities of current MLLMs through three key innovations: an efficient multi-span temporal grounding algorithm applied to low-dimension temporal features projected…
-
21 Dec 2023 2 repositories listed Syntology ran 8 of 10 samples · 2 unverifiedWe introduce LingoQA, a novel dataset and benchmark for visual question answering in autonomous driving.
Syntology lines on 24 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections