Datasets › OVBench

OVBench

Introduced by Zhenpeng Huang et al. in Online Video Understanding: OVBench and VideoChat-Online31 Dec 2024 archive 2025-07-28

OVBench is a benchmark tailored for real-time video understanding:

  • Memory, Perception, and Prediction of Temporal Contexts: Questions are framed to reference the present state of entities, requiring models to memorize/perceive/predict past/present/future temporal contexts over time.
  • Dynamic Spatio-temporal Interaction: The benchmark demands precise real-time interactions with video content, where actions, objects, and events must be understood in the context of their spatial and temporal relationships.
  • Contextual Awareness at Specific Moments: Real-time questions are contextual, changing based on the specific timestamp they are asked, requiring a deep understanding of how temporal context evolves.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Video Question Answering OVBench Seed1.5-VL AVG 60.0 Seed1.5-VL Technical Report — 16 Compare

Papers archive 2025-07-28

14 shown of 14 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 14. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Seed1.5-VL Technical Report 0 1 11 May 2025 not harvested
Online Video Understanding: OVBench and VideoChat-Online 1 1 31 Dec 2024 not harvested
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling 1 2 6 Dec 2024 ran 1 of 9 samples (8 unverified)
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution 8 1 18 Sep 2024 ran 8 of 12 samples (4 unverified)
LLaVA-OneVision: Easy Visual Task Transfer 2 1 6 Aug 2024 not harvested
Long Context Transfer from Language to Vision 2 1 24 Jun 2024 ran 5 of 5 samples (0 unverified; 5 pointer-only for licence)
VideoLLM-online: Online Video Large Language Model for Streaming Video 0 1 17 Jun 2024 not harvested
Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams 1 1 12 Jun 2024 not harvested
LITA: Language Instructed Temporal-Localization Assistant 1 1 27 Mar 2024 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context 1 1 8 Mar 2024 not harvested
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding 2 1 4 Dec 2023 ran 7 of 11 samples (4 unverified)
VTimeLLM: Empower LLM to Grasp Video Moments 1 1 30 Nov 2023 ran 5 of 11 samples (6 unverified; 11 pointer-only for licence)
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models 2 1 28 Nov 2023 ran 3 of 4 samples (1 unverified)
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding 1 1 31 Jul 2023 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • OVBench

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections