Papers › Relational Reasoning Over Spatial-Temporal Graphs for Video Summarization

Relational Reasoning Over Spatial-Temporal Graphs for Video Summarization

6 Apr 2022IEEE Transactions on Image Processing 2022 4archive 2025-07-28

Wencheng Zhu, Yucheng Han, Jiwen Lu, Jie zhou

In this paper, we propose a dynamic graph modeling approach to learn spatial-temporal representations for video summarization. Most existing video summarization methods extract image-level features with ImageNet pre-trained deep models. Differently, our method exploits object-level and relation-level information to capture spatial-temporal dependencies. Specifically, our method builds spatial graphs on the detected object proposals. Then, we construct a temporal graph by using the aggregated representations of spatial graphs. Afterward, we perform relational reasoning over spatial and temporal graphs with graph convolutional networks and extract spatial-temporal representations for importance score prediction and key shot selection. To eliminate relation clutters caused by densely connected nodes, we further design a self-attention edge pooling module, which disregards meaningless relations of graphs. We conduct extensive experiments on two popular benchmarks, including the SumMe and TVSum datasets. Experimental results demonstrate that the proposed method achieves superior performance against state-of-the-art video summarization methods.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Graph ClassificationRelational ReasoningSupervised Video SummarizationVideo Summarization

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Graph Classification NCI1 SAEPool_g Accuracy 74.48% #53 of 69 Archive leaderboard report
Graph Classification NCI109 SAEPool_h Accuracy 75.85 #25 of 38 Archive leaderboard report
Graph Classification PROTEINS SAEPool Accuracy 80.36% #9 of 103 Archive leaderboard report
Supervised Video Summarization SumMe RR-STG F1-score (Augmented) 54.8 #7 of 21 Archive leaderboard report
Supervised Video Summarization SumMe RR-STG F1-score (Canonical) 53.4 #7 of 21 Archive leaderboard report
Supervised Video Summarization SumMe RR-STG Kendall's Tau 0.211 #7 of 21 Archive leaderboard report
Supervised Video Summarization SumMe RR-STG Spearman's Rho 0.234 #7 of 21 Archive leaderboard report
Supervised Video Summarization TvSum RR-STG F1-score (Augmented) 63.6 #7 of 21 Archive leaderboard report
Supervised Video Summarization TvSum RR-STG F1-score (Canonical) 63.0 #7 of 21 Archive leaderboard report
Supervised Video Summarization TvSum RR-STG Kendall's Tau 0.162 #7 of 21 Archive leaderboard report
Supervised Video Summarization TvSum RR-STG Spearman's Rho 0.212 #7 of 21 Archive leaderboard report
Video Summarization SumMe RR-STG F1-score (Augmented) 55.3 #2 of 6 Archive leaderboard report
Video Summarization SumMe RR-STG F1-score (Canonical) 54.5 #2 of 6 Archive leaderboard report
Video Summarization TvSum RR-STG F1-score (Augmented) 63.6 #1 of 6 Archive leaderboard report
Video Summarization TvSum RR-STG F1-score (Canonical) 63.0 #1 of 6 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections