Papers › Skip-Clip: Self-Supervised Spatiotemporal Representation Learning by Future Clip Order Ranking

Skip-Clip: Self-Supervised Spatiotemporal Representation Learning by Future Clip Order Ranking

28 Oct 2019arXiv:1910.12770archive 2025-07-28

Alaaeldin El-Nouby, Shuangfei Zhai, Graham W. Taylor, Joshua M. Susskind

Deep neural networks require collecting and annotating large amounts of data to train successfully. In order to alleviate the annotation bottleneck, we propose a novel self-supervised representation learning approach for spatiotemporal features extracted from videos. We introduce Skip-Clip, a method that utilizes temporal coherence in videos, by training a deep model for future clip order ranking conditioned on a context clip as a surrogate objective for video future prediction. We show that features learned using our method are generalizable and transfer strongly to downstream tasks. For action recognition on the UCF101 dataset, we obtain 51.8% improvement over random initialization and outperform models initialized using inflated ImageNet parameters. Skip-Clip also achieves results competitive with state-of-the-art self-supervision methods.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionFuture predictionRepresentation LearningSelf-Supervised Action Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Self-Supervised Action Recognition UCF101 Skip-Clip (3D ResNet-18) 3-fold Accuracy 64.4 #44 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 Skip-Clip (3D ResNet-18) Frozen false #44 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 Skip-Clip (3D ResNet-18) Pre-Training Dataset UCF101 #44 of 53 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

1x1 ConvolutionAverage PoolingBatch NormalizationBottleneck Residual BlockConvolutionGlobal Average PoolingKaiming InitializationMax PoolingReLUResidual BlockResidual Connection

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections