Papers › Self-Supervised Spatiotemporal Feature Learning via Video Rotation Prediction

Self-Supervised Spatiotemporal Feature Learning via Video Rotation Prediction

28 Nov 2018arXiv:1811.11387archive 2025-07-28

Longlong Jing, Xiaodong Yang, Jingen Liu, YingLi Tian

The success of deep neural networks generally requires a vast amount of training data to be labeled, which is expensive and unfeasible in scale, especially for video collections. To alleviate this problem, in this paper, we propose 3DRotNet: a fully self-supervised approach to learn spatiotemporal features from unlabeled videos. A set of rotations are applied to all videos, and a pretext task is defined as prediction of these rotations. When accomplishing this task, 3DRotNet is actually trained to understand the semantic concepts and motions in videos. In other words, it learns a spatiotemporal video representation, which can be transferred to improve video understanding tasks in small datasets. Our extensive experiments successfully demonstrate the effectiveness of the proposed framework on action recognition, leading to significant improvements over the state-of-the-art self-supervised methods. With the self-supervised pre-trained 3DRotNet from large datasets, the recognition accuracy is boosted up by 20.4% on UCF101 and 16.7% on HMDB51 respectively, compared to the models trained from scratch.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionPredictionSelf-Supervised Action RecognitionTemporal Action LocalizationVideo Understanding

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Self-Supervised Action Recognition HMDB51 3D RotNet (3D ResNet-18) Frozen false #42 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 3D RotNet (3D ResNet-18) Pre-Training Dataset Kinetics400 #42 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 3D RotNet (3D ResNet-18) Top-1 Accuracy 33.7 #42 of 48 Archive leaderboard report
Self-Supervised Action Recognition UCF101 3D RotNet (3D ResNet-18) 3-fold Accuracy 62.9 #45 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 3D RotNet (3D ResNet-18) Frozen false #45 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 3D RotNet (3D ResNet-18) Pre-Training Dataset Kinetics400 #45 of 53 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

1x1 ConvolutionAverage PoolingBatch NormalizationBottleneck Residual BlockConvolutionGlobal Average PoolingKaiming InitializationMax PoolingReLUResidual BlockResidual Connection

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections