Browse State-of-the-Art › Zero-Shot Video Question Answer
Zero-Shot Video Question Answer
73 papers with code · 17 benchmarks · 18 datasets archive 2025-07-28
This task present the results of Zeroshot Question Answer results on TGIF-QA dataset for LLM powered Video Conversational Models.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
17 leaderboard tables shown for this task, 17 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 17 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
18 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 73 papers with code (85 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
18 Sep 2024 8 repositories listed Syntology ran 8 of 12 samples · 4 unverifiedWe present the Qwen2-VL Series, an advanced upgrade of the previous Qwen-VL models that redefines the conventional predetermined-resolution approach in visual processing.
-
16 Nov 2023 6 repositories listed Syntology ran 4 of 7 samples · 3 unverified · 1 pointer-only (licence)In this work, we unify visual representation into the language feature space to advance the foundational LLM towards a unified LVLM.
-
10 Oct 2023 6 repositories listed Syntology ran 9 of 11 samples · 2 unverified · 1 pointer-only (licence)We introduce Mistral 7B v0.
-
29 Apr 2022 5 repositories listed Syntology ran 18 of 24 samples · 6 unverified · 7 pointer-only (licence)Building models that can be rapidly adapted to novel tasks using only a handful of annotated examples is an open challenge for multimodal machine learning research.
-
14 Nov 2023 4 repositories listedLarge language models have demonstrated impressive universal capabilities across a wide range of open-ended tasks and have extended their utility to encompass multimodal conversations.
-
5 Jun 2023 4 repositories listed Syntology ran 18 of 25 samples · 7 unverified · 9 pointer-only (licence)We present Video-LLaMA a multi-modal framework that empowers Large Language Models (LLMs) with the capability of understanding both visual and auditory content in the video.
-
10 Jul 2024 3 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 4 pointer-only (licence)To this end, we introduce LLaVA-NeXT-Interleave, which simultaneously tackles Multi-image, Multi-frame (video), Multi-view (3D), and Multi-patch (single-image) scenarios in LMMs.
-
11 Jun 2024 3 repositories listed Syntology ran 7 of 17 samples · 10 unverifiedIn this paper, we present the VideoLLaMA 2, a set of Video Large Language Models (Video-LLMs) designed to enhance spatial-temporal modeling and audio understanding in video and audio-oriented tasks.
-
12 Dec 2023 3 repositories listedVisual language models (VLMs) rapidly progressed with the recent success of large language models.
-
28 Nov 2023 3 repositories listed Syntology ran 7 of 10 samples · 3 unverifiedWith the rapid development of Multi-modal Large Language Models (MLLMs), a number of diagnostic benchmarks have recently emerged to evaluate the comprehension capabilities of these models.
-
28 Apr 2023 3 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)This strategy effectively alleviates the interference between the two tasks of image-text alignment and instruction following and achieves strong multi-modal reasoning with only a small-scale image-text and instruction…
-
16 Jun 2022 3 repositories listed Syntology ran 14 of 34 samples · 20 unverified · 1 pointer-only (licence)Manual annotation of question and answers for videos, however, is tedious and prohibits scalability.
-
6 Aug 2024 2 repositories listedWe present LLaVA-OneVision, a family of open large multimodal models (LMMs) developed by consolidating our insights into data, models, and visual representations in the LLaVA-NeXT blog series.
-
3 Aug 2024 2 repositories listed Syntology ran 9 of 14 samples · 5 unverifiedThe recent surge of Multimodal Large Language Models (MLLMs) has fundamentally reshaped the landscape of AI research and industry, shedding light on a promising path toward the next AI milestone.
-
24 Jun 2024 2 repositories listed Syntology ran 5 of 5 samples · 0 unverified · 5 pointer-only (licence)By simply extrapolating the context length of the language backbone, we enable LMMs to comprehend orders of magnitude more visual tokens without any video training.
-
4 Apr 2024 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedThis paper introduces MiniGPT4-Video, a multimodal Large Language Model (LLM) designed specifically for video understanding.
-
22 Mar 2024 2 repositories listedWe introduce InternVideo2, a new family of video foundation models (ViFM) that achieve the state-of-the-art results in video recognition, video-text tasks, and video-centric dialogue.
-
15 Mar 2024 2 repositories listedLong-form video understanding represents a significant challenge within computer vision, demanding a model capable of reasoning over long multi-modal sequences.
-
20 Feb 2024 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 1 pointer-only (licence)We utilize a curriculum learning training scheme to learn the hierarchical structure of videos, starting from clip-level captions describing atomic actions, then focusing on segment-level descriptions, and concluding…
-
4 Dec 2023 2 repositories listed Syntology ran 7 of 11 samples · 4 unverifiedThis work proposes TimeChat, a time-sensitive multimodal large language model specifically designed for long video understanding.
-
28 Nov 2023 2 repositories listed Syntology ran 3 of 4 samples · 1 unverifiedCurrent VLMs, while proficient in tasks like image captioning and visual question answering, face computational burdens when processing long videos due to the excessive visual tokens.
-
8 Jun 2023 2 repositories listedConversation agents fueled by Large Language Models (LLMs) are providing a new way to interact with visual data.
-
6 Dec 2022 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedSpecifically, InternVideo efficiently explores masked video modeling and video-language contrastive learning as the pretraining objectives, and selectively coordinates video representations of these two complementary…
-
26 Jul 2019 2 repositories listedSecond, all baggage images are captured by specially-designed multi-view camera system to handle pose variation and occlusion, in order to obtain the 3D information of baggage surface as complete as possible.
-
14 Apr 2017 2 repositories listedIn this paper, we focus on extending VQA to the video domain and contribute to the literature in three important ways.
-
25 Apr 2025 1 repository listedVideo Question Answering (VQA) inherently relies on multimodal reasoning, integrating visual, temporal, and linguistic cues to achieve a deeper understanding of video content.
-
26 Mar 2025 1 repository listedIn this framework, Thinker functions as a large language model tasked with text generation, while Talker is a dual-track autoregressive model that directly utilizes the hidden representations from the Thinker to produce…
-
20 Mar 2025 1 repository listed Syntology ran 3 of 8 samples · 5 unverifiedVideo question answering (VideoQA) enables machines to extract and comprehend key information from videos through natural language interaction, which is a critical step towards achieving intelligence.
-
17 Mar 2025 1 repository listed Syntology ran 3 of 4 samples · 1 unverifiedVideos, with their unique temporal dimension, demand precise grounded understanding, where answers are directly linked to visual, interpretable evidence.
-
12 Mar 2025 1 repository listedThe self-attention mechanism provides a general solution for sequence modeling, but it has a prohibitive cost when applied to a massive number of spatiotemporal tokens in long videos.
Syntology lines on 19 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections