Browse State-of-the-Art › Video-based Generative Performance Benchmarking (Consistency)
Video-based Generative Performance Benchmarking (Consistency)
15 papers with code · 1 benchmark · 1 dataset archive 2025-07-28
The benchmark evaluates a generative Video Conversational Model with respect to Consistency.
We curate a test set based on the ActivityNet-200 dataset, featuring videos with rich, dense descriptive captions and associated question-answer pairs from human annotations. We develop an evaluation pipeline using the GPT-3.5 model that assigns a relative score to the generated predictions on a scale of 1-5.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| VideoInstruct (18 rows) | PPLLaVA-7B | PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance | code | Syntology ran 2 of 9 samples · 7 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
15 shown of 15 papers with code (15 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
14 Nov 2023 4 repositories listedLarge language models have demonstrated impressive universal capabilities across a wide range of open-ended tasks and have extended their utility to encompass multimodal conversations.
-
5 Jun 2023 4 repositories listed Syntology ran 18 of 25 samples · 7 unverified · 9 pointer-only (licence)We present Video-LLaMA a multi-modal framework that empowers Large Language Models (LLMs) with the capability of understanding both visual and auditory content in the video.
-
28 Nov 2023 3 repositories listed Syntology ran 7 of 10 samples · 3 unverifiedWith the rapid development of Multi-modal Large Language Models (MLLMs), a number of diagnostic benchmarks have recently emerged to evaluate the comprehension capabilities of these models.
-
28 Apr 2023 3 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)This strategy effectively alleviates the interference between the two tasks of image-text alignment and instruction following and achieves strong multi-modal reasoning with only a small-scale image-text and instruction…
-
4 Apr 2024 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedThis paper introduces MiniGPT4-Video, a multimodal Large Language Model (LLM) designed specifically for video understanding.
-
8 Jun 2023 2 repositories listedConversation agents fueled by Large Language Models (LLMs) are providing a new way to interact with visual data.
-
17 Nov 2024 1 repository listed Syntology ran 5 of 11 samples · 6 unverifiedFor video understanding tasks, training-based video LLMs are difficult to build due to the scarcity of high-quality, curated video-text paired data.
-
4 Nov 2024 1 repository listed Syntology ran 2 of 9 samples · 7 unverified · 3 pointer-only (licence)In this paper, we identify the key issue as the redundant content in videos.
-
22 Jul 2024 1 repository listed Syntology ran 4 of 6 samples · 2 unverified · 6 pointer-only (licence)As a result, this design allows us to adequately capture both spatial and temporal features that are beneficial for detailed video understanding.
-
13 Jun 2024 1 repository listed Syntology ran 6 of 8 samples · 2 unverified · 8 pointer-only (licence)Building on the advances of language models, Large Multimodal Models (LMMs) have contributed significant improvements in video understanding.
-
25 Apr 2024 1 repository listed Syntology ran 0 of 2 samples · 2 unverified · 2 pointer-only (licence)PLLaVA achieves new state-of-the-art performance on modern benchmark datasets for both video question-answer and captioning tasks.
-
30 Nov 2023 1 repository listed Syntology ran 5 of 11 samples · 6 unverified · 11 pointer-only (licence)Large language models (LLMs) have shown remarkable text understanding capabilities, which have been extended as Video LLMs to handle video data for comprehending visual details.
-
27 Sep 2023 1 repository listed Syntology ran 1 of 4 samples · 3 unverifiedWithout bells and whistles, BT-Adapter achieves (1) state-of-the-art zero-shot results on various video tasks using thousands of fewer GPU hours.
-
31 Jul 2023 1 repository listedRecently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks.
-
10 May 2023 1 repository listedIn this paper, we initiate an attempt of developing an end-to-end chat-centric video understanding system, coined as VideoChat.
Syntology lines on 11 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections