Papers › Progressive Video Summarization via Multimodal Self-supervised Learning

Progressive Video Summarization via Multimodal Self-supervised Learning

7 Jan 2022arXiv:2201.02494archive 2025-07-28

Li Haopeng, Ke Qiuhong, Gong Mingming, Tom Drummond

Modern video summarization methods are based on deep neural networks that require a large amount of annotated data for training. However, existing datasets for video summarization are small-scale, easily leading to over-fitting of the deep models. Considering that the annotation of large-scale datasets is time-consuming, we propose a multimodal self-supervised learning framework to obtain semantic representations of videos, which benefits the video summarization task. Specifically, the self-supervised learning is conducted by exploring the semantic consistency between the videos and text in both coarse-grained and fine-grained fashions, as well as recovering masked frames in the videos. The multimodal framework is trained on a newly-collected dataset that consists of video-text pairs. Additionally, we introduce a progressive video summarization method, where the important content in a video is pinpointed progressively to generate better summaries. Extensive experiments have proved the effectiveness and superiority of our method in rank correlation coefficients and F-score.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

HopLee6/SSPVS-PyTorch officialmentioned on GitHubpytorch report
thswodnjs3/CSTA mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Self-Supervised LearningSupervised Video SummarizationVideo ClassificationVideo Summarization

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Supervised Video Summarization SumMe SSPVS(+Text) F1-score (Canonical) 50.7 #10 of 21 Archive leaderboard report
Supervised Video Summarization SumMe SSPVS(+Text) Kendall's Tau 0.192 #10 of 21 Archive leaderboard report
Supervised Video Summarization SumMe SSPVS(+Text) Spearman's Rho 0.257 #10 of 21 Archive leaderboard report
Supervised Video Summarization SumMe SSPVS F1-score (Augmented) 50.4 #13 of 21 Archive leaderboard report
Supervised Video Summarization SumMe SSPVS F1-score (Canonical) 48.7 #13 of 21 Archive leaderboard report
Supervised Video Summarization SumMe SSPVS Kendall's Tau 0.178 #13 of 21 Archive leaderboard report
Supervised Video Summarization SumMe SSPVS Spearman's Rho 0.240 #13 of 21 Archive leaderboard report
Supervised Video Summarization TvSum SSPVS(+Text) F1-score (Canonical) 60.4 #15 of 21 Archive leaderboard report
Supervised Video Summarization TvSum SSPVS(+Text) Kendall's Tau 0.181 #15 of 21 Archive leaderboard report
Supervised Video Summarization TvSum SSPVS(+Text) Spearman's Rho 0.238 #15 of 21 Archive leaderboard report
Supervised Video Summarization TvSum SSPVS F1-score (Augmented) 61.8 #16 of 21 Archive leaderboard report
Supervised Video Summarization TvSum SSPVS F1-score (Canonical) 60.3 #16 of 21 Archive leaderboard report
Supervised Video Summarization TvSum SSPVS Kendall's Tau 0.177 #16 of 21 Archive leaderboard report
Supervised Video Summarization TvSum SSPVS Spearman's Rho 0.233 #16 of 21 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections