Datasets › TVBench

TVBench

Introduced by Daniel Cores et al. in TVBench: Redesigning Video-Language Evaluation10 Oct 2024 archive 2025-07-28

TVBench is a new benchmark specifically created to evaluate temporal understanding in video QA. We identified three main issues in existing datasets: (i) static information from single frames is often sufficient to solve the tasks (ii) the text of the questions and candidate answers is overly informative, allowing models to answer correctly without relying on any visual input (iii) world knowledge alone can answer many of the questions, making the benchmarks a test of knowledge replication rather than visual reasoning. In addition, we found that open-ended question-answering benchmarks for video understanding suffer from similar issues while the automatic evaluation process with LLMs is unreliable, making it an unsuitable alternative.

We defined 10 temporally challenging tasks that either require repetition counting (Action Count), properties about moving objects (Object Shuffle, Object Count, Moving Direction), temporal localization (Action Localization, Unexpected Action), temporal sequential ordering (Action Sequence, Scene Transition, Egocentric Sequence) and distinguishing between temporally hard Action Antonyms such as "Standing up" and "Sitting down".

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Video Question Answering TVBench Seed1.5-VL thinking Average Accuracy 63.6 Seed1.5-VL Technical Report — 28 Compare

Papers archive 2025-07-28

18 shown of 18 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 22. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning 1 1 11 Jun 2025 ran 0 of 8 samples (8 unverified; 8 pointer-only for licence)
Seed1.5-VL Technical Report 0 2 11 May 2025 not harvested
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding 1 3 17 Apr 2025 not harvested
Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization 0 1 16 Apr 2025 not harvested
Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding 1 1 14 Jan 2025 ran 4 of 4 samples (0 unverified; 4 pointer-only for licence)
GPT-4o System Card 0 1 25 Oct 2024 not harvested
Aria: An Open Multimodal Native Mixture-of-Experts Model 1 1 8 Oct 2024 ran 1 of 1 samples (0 unverified)
Video Instruction Tuning With Synthetic Data 0 2 3 Oct 2024 not harvested
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution 8 2 18 Sep 2024 ran 8 of 12 samples (4 unverified)
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 1 1 9 Aug 2024 not harvested
InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output 1 1 3 Jul 2024 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
Tarsier: Recipes for Training and Evaluating Large Video Description Models 1 2 30 Jun 2024 ran 2 of 2 samples (0 unverified)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding 1 1 13 Jun 2024 ran 6 of 8 samples (2 unverified; 8 pointer-only for licence)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs 3 3 11 Jun 2024 ran 7 of 17 samples (10 unverified)
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning 1 3 25 Apr 2024 ran 0 of 2 samples (2 unverified; 2 pointer-only for licence)
ST-LLM: Large Language Models Are Effective Temporal Learners 1 1 30 Mar 2024 ran 7 of 11 samples (4 unverified; 3 pointer-only for licence)
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context 1 1 8 Mar 2024 not harvested
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark 3 1 28 Nov 2023 ran 7 of 10 samples (3 unverified)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

cc-by-4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • TVBench

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections