Papers › Self-Supervised Spatiotemporal Learning via Video Clip Order Prediction

Self-Supervised Spatiotemporal Learning via Video Clip Order Prediction

1 Jun 2019CVPR 2019 6archive 2025-07-28

Dejing Xu, Jun Xiao, Zhou Zhao, Jian Shao, Di Xie, Yueting Zhuang

We propose a self-supervised spatiotemporal learning technique which leverages the chronological order of videos. Our method can learn the spatiotemporal representation of the video by predicting the order of shuffled clips from the video. The category of the video is not required, which gives our technique the potential to take advantage of infinite unannotated videos. There exist related works which use frames, while compared to frames, clips are more consistent with the video dynamics. Clips can help to reduce the uncertainty of orders and are more appropriate to learn a video representation. The 3D convolutional neural networks are utilized to extract features for clips, and these features are processed to predict the actual order. The learned representations are evaluated via nearest neighbor retrieval experiments. We also use the learned networks as the pre-trained models and finetune them on the action recognition task. Three types of 3D convolutional neural networks are tested in experiments, and we gain large improvements compared to existing self-supervised methods.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionRetrievalSelf-Supervised Action RecognitionTemporal Action Localization

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Self-Supervised Action Recognition HMDB51 Video Clip Ordering (R3D) Frozen false #45 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 Video Clip Ordering (R3D) Pre-Training Dataset UCF101 #45 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 Video Clip Ordering (R3D) Top-1 Accuracy 29.5 #45 of 48 Archive leaderboard report
Self-Supervised Action Recognition UCF101 Video Clip Ordering (R3D) 3-fold Accuracy 64.9 #43 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 Video Clip Ordering (R3D) Frozen false #43 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 Video Clip Ordering (R3D) Pre-Training Dataset UCF101 #43 of 53 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections