Browse State-of-the-Art › Video Summarization
Video Summarization
89 papers with code · 6 benchmarks · 16 datasets archive 2025-07-28
Video Summarization aims to generate a short synopsis that summarizes the video content by selecting its most informative and important parts. The produced summary is usually composed of a set of representative video frames (a.k.a. video key-frames), or video fragments (a.k.a. video key-fragments) that have been stitched in chronological order to form a shorter video. The former type of a video summary is known as video storyboard, and the latter type is known as video skim.
Source: Video Summarization Using Deep Neural Networks: A Survey
Image credit: iJRASET
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
6 leaderboard tables shown for this task, 6 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| SumMe (6 rows) | PGL-SUM | Combining Global and Local Attention with Positional Encoding for... | code | — | Compare |
| TvSum (6 rows) | RR-STG | Relational Reasoning Over Spatial-Temporal Graphs for Video Summarization | — | — | Compare |
| Query-Focused Video Summarization Dataset (2 rows) | EgoVLPv2 | EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in... | code | — | Compare |
| Shot2Story20K (2 rows) | Shotluck-Holmes (3.1B) | Shotluck Holmes: A Family of Efficient Small-Scale Large Language... | code | — | Compare |
| Mr. HiSum (1 row) | PGL-SUM | Mr. HiSum: A Large-scale Dataset for Video Highlight Detection and... | — | — | Compare |
| videoxum (1 row) | VTSUM-BLIP | VideoXum: Cross-modal Visual and Textural Summarization of Videos | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
16 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
2 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 89 papers with code (280 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
29 Dec 2017 6 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedVideo summarization aims to facilitate large-scale video browsing by producing short, concise summaries that are diverse and representative of original videos.
-
5 Dec 2018 5 repositories listed Syntology ran 1 of 15 samples · 14 unverifiedIn this work we propose a novel method for supervised, keyshots based video summarization by applying a conceptually simple and computationally efficient soft, self-attention mechanism.
-
26 Apr 2021 4 repositories listedTraditional video summarization methods generate fixed video representations regardless of user interest.
-
23 Apr 2021 3 repositories listedThe proposed architecture utilizes an attention mechanism before fusing motion features and features representing the (static) visual content, i.
-
10 Oct 2019 3 repositories listedVideo is one of the robust sources of information and the consumption of online and offline videos has reached an unprecedented level in the last few years.
-
7 Jun 2023 2 repositories listedTo address these challenges and provide a comprehensive dataset for this new direction, we have meticulously curated the \textbf{MMSum} dataset.
-
13 Mar 2023 2 repositories listedThe goal of multimodal summarization is to extract the most important information from different modalities to form output summaries.
-
3 Jun 2022 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 2 pointer-only (licence)Video-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention.
-
7 Jan 2022 2 repositories listedConsidering that the annotation of large-scale datasets is time-consuming, we propose a multimodal self-supervised learning framework to obtain semantic representations of videos, which benefits the video summarization…
-
27 Dec 2021 2 repositories listedVideo summarization aims to automatically generate a summary (storyboard or video skim) of a video, which can facilitate large-scale video retrieval and browsing.
-
31 Jan 2020 2 repositories listedThis paper addresses the task of query-focused video summarization, which takes user's query and a long video as inputs and aims to generate a query-focused video summary.
-
27 Mar 2019 2 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedVideo summarization is a technique to create a short skim of the original video while preserving the main stories/content.
-
29 Nov 2018 2 repositories listedIn our algorithm, at each iteration, the maximum information from the structure of the data is captured by one selected sample, and the captured information is neglected in the next iterations by projection on the…
-
28 Sep 2016 2 repositories listedFor this, we design a deep neural network that maps videos as well as descriptions to a common semantic space and jointly trained it with associated pairs of videos and descriptions.
-
6 May 2025 1 repository listedIn this work, we introduce the task of script-driven video summarization, which aims to produce a summary of the full-length video by selecting the parts that are most relevant to a user-provided script outlining the…
-
6 Mar 2025 1 repository listedWe present Vinci, a vision-language system designed to provide real-time, comprehensive AI assistance on portable devices.
-
26 Feb 2025 1 repository listedIn this work, we present a novel unsupervised scheme named SegSum, designed for video summarization through the creation of video skims.
-
12 Feb 2025 1 repository listedTransforming recorded videos into concise and accurate textual summaries is a growing challenge in multimodal learning.
-
18 Dec 2024 1 repository listedCurrent approaches to video understanding with LLMs often rely on pretrained video encoders to extract spatiotemporal features and text encoders to capture semantic meaning.
-
12 Dec 2024 1 repository listed Syntology ran 0 of 11 samples · 11 unverifiedThe demand for producing short-form videos for sharing on social media platforms has experienced significant growth in recent times.
-
17 Oct 2024 1 repository listedThis paper introduces an approach for query-focused video summarization, aiming to align video summaries closely with user queries.
-
4 Oct 2024 1 repository listedThese models achieved state-of-the-art correlation scores on important benchmark datasets.
-
22 Sep 2024 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 2 pointer-only (licence)The VST module is trained by instruction fine-tuning, where two optimizing strategies are offered.
-
24 Jun 2024 1 repository listedWith the surge in the amount of video data, video summarization techniques, including visual-modal(VM) and textual-modal(TM) summarization, are attracting more and more attention.
-
5 Jun 2024 1 repository listedIn this paper we introduce a new dataset for 360-degree video summarization: the transformation of 360-degree video content to concise 2D-video summaries that can be consumed via traditional devices, such as TV sets and…
-
31 May 2024 1 repository listedVideo is an increasingly prominent and information-dense medium, yet it poses substantial challenges for language models.
-
22 May 2024 1 repository listed Syntology ran 13 of 17 samples · 4 unverifiedVideo Temporal Grounding (VTG) strives to accurately pinpoint event timestamps in a specific video using linguistic queries, significantly impacting downstream tasks like video browsing and editing.
-
20 May 2024 1 repository listed Syntology ran 7 of 8 samples · 1 unverifiedVideo summarization aims to generate a concise representation of a video, capturing its essential content and key moments while reducing its overall length.
-
16 May 2024 1 repository listedThe findings of the conducted quantitative and qualitative evaluations demonstrate the ability of our framework to spot the most and least influential fragments and visual objects of the video for the summarizer, and to…
-
14 Apr 2024 1 repository listedWe propose a graph-based representation learning framework for video summarization.
Syntology lines on 8 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections