Browse State-of-the-Art › Highlight Detection
Highlight Detection
41 papers with code · 4 benchmarks · 2 datasets archive 2025-07-28
https://youtu.be/pJ0auP7dbcY?si=vSiZevfJ57YUKC2q
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
4 leaderboard tables shown for this task, 4 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| QVHighlights (21 rows) | SG-DETR (w/ PT) | Saliency-Guided DETR for Moment Retrieval and Highlight Detection | code | — | Compare |
| TvSum (7 rows) | FlashVTG | FlashVTG: Feature Layering and Adaptive Score Handling Network for... | code | Syntology ran 2 of 13 samples · 11 unverified | Compare |
| YouTube Highlights (7 rows) | SG-DETR (w/ PT) | Saliency-Guided DETR for Moment Retrieval and Highlight Detection | code | — | Compare |
| arabiska (1 row) | Kenan Kanan | 100,000 Podcasts: A Spoken English Document Corpus | — | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 41 papers with code (78 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
20 Jul 2021 4 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 1 pointer-only (licence)Each video in the dataset is annotated with: (1) a human-written free-form NL query, (2) relevant moments in the video w.
-
23 Mar 2022 3 repositories listedFinding relevant moments and highlights in videos according to natural language queries is a natural and highly valuable common need in the current video content explosion era.
-
4 Dec 2023 2 repositories listed Syntology ran 7 of 11 samples · 4 unverifiedThis work proposes TimeChat, a time-sensitive multimodal large language model specifically designed for long video understanding.
-
15 Nov 2023 2 repositories listed Syntology ran 2 of 11 samples · 9 unverified · 11 pointer-only (licence)Dummy tokens conditioned by text query take portions of the attention weights, preventing irrelevant video clips from being represented by the text query.
-
26 May 2025 1 repository listedPodcasts have become daily companions for half a billion users.
-
2 May 2025 1 repository listedWe propose TEMPURA (Temporal Event Masked Prediction and Understanding for Reasoning in Action), a two-stage training framework that enhances video temporal understanding.
-
18 Jan 2025 1 repository listedExisting models usually first use contrastive learning methods to align video and text features, then fuse and extract multimodal information, and finally use a Transformer Decoder to decode multimodal information.
-
5 Jan 2025 1 repository listed Syntology ran 8 of 10 samples · 2 unverified · 10 pointer-only (licence)In this paper, we present a novel Video Context-aware Keyword Attention module that overcomes this limitation by capturing keyword variation within the context of the entire video.
-
18 Dec 2024 1 repository listed Syntology ran 2 of 13 samples · 11 unverified · 13 pointer-only (licence)For short-moment retrieval, FlashVTG increases mAP to 125% of previous SOTA performance.
-
12 Dec 2024 1 repository listed Syntology ran 0 of 11 samples · 11 unverifiedThe demand for producing short-form videos for sharing on social media platforms has experienced significant growth in recent times.
-
2 Dec 2024 1 repository listedVideo Highlight Detection and Moment Retrieval (HD/MR) are essential in video analysis.
-
27 Nov 2024 1 repository listed Syntology ran 0 of 2 samples · 2 unverifiedWe construct MMDuetIT, a video-text training dataset designed to adapt VideoLLMs to video-text duet interaction format.
-
15 Nov 2024 1 repository listed Syntology ran 1 of 11 samples · 10 unverifiedVideo Large Language Models (Vid-LLMs) have made remarkable advancements in comprehending video content for QA dialogue.
-
25 Oct 2024 1 repository listed Syntology ran 3 of 6 samples · 3 unverifiedThis paper proposes TimeSuite, a collection of new designs to adapt the existing short-form video MLLMs for long video understanding, including a simple yet efficient framework to process long video sequence, a…
-
2 Oct 2024 1 repository listedCombined with the introduced Saliency-Guided Cross Attention mechanism and a hybrid DETR architecture, our approach significantly enhances performance in both moment retrieval and highlight detection tasks.
-
6 Aug 2024 1 repository listed Syntology ran 5 of 11 samples · 6 unverifiedLighthouse addresses these issues by implementing a unified reproducible codebase that includes six models, three features, and five datasets.
-
21 Jul 2024 1 repository listed Syntology ran 5 of 11 samples · 6 unverifiedThrough a feasibility study, we demonstrate that LLM encoders effectively refine inter-concept relations in multimodal embeddings, even without being trained on textual embeddings.
-
22 May 2024 1 repository listed Syntology ran 13 of 17 samples · 4 unverifiedVideo Temporal Grounding (VTG) strives to accurately pinpoint event timestamps in a specific video using linguistic queries, significantly impacting downstream tasks like video browsing and editing.
-
14 Apr 2024 1 repository listed Syntology ran 6 of 10 samples · 4 unverified · 10 pointer-only (licence)Video moment retrieval and highlight detection are two highly valuable tasks in video understanding, but until recently they have been jointly studied.
-
2 Apr 2024 1 repository listedVideo temporal grounding (VTG) is a fine-grained video understanding problem that aims to ground relevant clips in untrimmed videos given natural language queries.
-
2 Apr 2024 1 repository listed Syntology ran 3 of 7 samples · 4 unverifiedMultimodal and large language models (LLMs) have revolutionized the utilization of open-world knowledge, unlocking novel potentials across various tasks and applications.
-
31 Mar 2024 1 repository listedVideo temporal grounding (VTG) is a fine-grained video understanding problem that aims to ground relevant clips in untrimmed videos given natural language queries.
-
4 Jan 2024 1 repository listed Syntology ran 11 of 17 samples · 6 unverified · 17 pointer-only (licence)Finally, a task cooperation module is constructed to refine the retrieval pipeline and the highlight score prediction process by utilizing the reciprocity between MR and HD.
-
28 Nov 2023 1 repository listed Syntology ran 9 of 12 samples · 3 unverifiedVideo Moment Retrieval (MR) and Highlight Detection (HD) have attracted significant attention due to the growing demand for video analysis.
-
28 Nov 2023 1 repository listedOnce trained, SHMGAN is able to generate specular-free images from a single RGB image as input; without requiring any additional external labels.
-
31 Jul 2023 1 repository listed Syntology ran 10 of 16 samples · 6 unverifiedMost methods in this direction develop taskspecific models that are trained with type-specific labels, such as moment retrieval (time interval) and highlight detection (worthiness curve), which limits their abilities to…
-
8 May 2023 1 repository listedVideo summarization has become an increasingly important task in the field of computer vision due to the vast amount of video content available on the internet.
-
29 Apr 2023 1 repository listedWith the increasing demand for video understanding, video moment and highlight detection (MHD) has emerged as a critical research topic.
-
26 Mar 2023 1 repository listedBased on existing efforts, this work has two observations: (1) For different annotators, labeling highlight has uncertainty, which leads to inaccurate and time-consuming annotations.
-
24 Mar 2023 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)As we observe the insignificant role of a given query in transformer architectures, our encoding module starts with cross-attention layers to explicitly inject the context of text query into video representation.
Syntology lines on 18 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections