Datasets › QVHighlights

QVHighlights (Query-based Video Highlights)

Introduced by Jie Lei et al. in QVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries20 Jul 2021 archive 2025-07-28

The Query-based Video Highlights (QVHighlights) dataset is a dataset for detecting customized moments and highlights from videos given natural language (NL). It consists of over 10,000 YouTube videos, covering a wide range of topics, from everyday activities and travel in lifestyle vlog videos to social and political activities in news videos. Each video in the dataset is annotated with: (1) a human-written free-form NL query, (2) relevant moments in the video w.r.t. the query, and (3) five-point scale saliency scores for all query-relevant clips.

Benchmarks archive 2025-07-28

All 4 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

26 shown of 26 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 41. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
LD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight Detection 1 1 18 Jan 2025 not harvested
Length-Aware DETR for Robust Moment Retrieval 1 1 30 Dec 2024 not harvested
FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding 1 2 18 Dec 2024 ran 2 of 13 samples (11 unverified; 13 pointer-only for licence)
VideoLights: Feature Refinement and Cross-Task Alignment Transformer for Joint Video Highlight Detection and Moment Retrieval 1 2 2 Dec 2024 not harvested
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval 1 1 21 Nov 2024 not harvested
Number it: Temporal Grounding Videos like Flipping Manga 1 1 15 Nov 2024 ran 1 of 11 samples (10 unverified)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning 1 1 25 Oct 2024 ran 3 of 6 samples (3 unverified)
Saliency-Guided DETR for Moment Retrieval and Highlight Detection 1 5 2 Oct 2024 not harvested
Prior Knowledge Integration via LLM Encoding and Pseudo Event Regulation for Video Moment Retrieval 1 3 21 Jul 2024 ran 5 of 11 samples (6 unverified)
Unleash the Potential of CLIP for Video Highlight Detection 1 1 2 Apr 2024 ran 3 of 7 samples (4 unverified)
R²-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding 1 2 31 Mar 2024 not harvested
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding 2 3 22 Mar 2024 not harvested
Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding 1 1 14 Mar 2024 ran 8 of 13 samples (5 unverified)
VTG-GPT: Tuning-Free Zero-Shot Video Temporal Grounding with GPT 1 1 4 Mar 2024 ran 9 of 10 samples (1 unverified)
BAM-DETR: Boundary-Aligned Moment Detection Transformer for Temporal Sentence Grounding in Videos 1 3 30 Nov 2023 ran 8 of 11 samples (3 unverified; 11 pointer-only for licence)
Bridging the Gap: A Unified Video Comprehension Framework for Moment Retrieval and Highlight Detection 1 2 28 Nov 2023 ran 9 of 12 samples (3 unverified)
Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding 2 4 15 Nov 2023 ran 2 of 11 samples (9 unverified; 11 pointer-only for licence)
DiffusionVMR: Diffusion Model for Joint Video Moment Retrieval and Highlight Detection 0 1 29 Aug 2023 not harvested
UnLoc: A Unified Framework for Video Localization Tasks 1 2 21 Aug 2023 not harvested
UniVTG: Towards Unified Video-Language Temporal Grounding 1 4 31 Jul 2023 ran 10 of 16 samples (6 unverified)
Background-aware Moment Detection for Video Moment Retrieval 1 1 5 Jun 2023 not harvested
Boundary-Denoising for Video Activity Localization 1 1 6 Apr 2023 not harvested
Query-Dependent Video Representation for Moment Retrieval and Highlight Detection 1 9 24 Mar 2023 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
UMT: Unified Multi-modal Transformers for Joint Video Moment Retrieval and Highlight Detection 3 5 23 Mar 2022 not harvested
Detecting Moments and Highlights in Videos via Natural Language Queries 1 1 1 Dec 2021 not harvested
QVHighlights: Detecting Moments and Highlights in Videos via Natural Language Queries 4 2 20 Jul 2021 ran 2 of 2 samples (0 unverified; 1 pointer-only for licence)

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Attribution-NonCommercial-ShareAlike 4.0 International

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • QVHighlights

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections