Papers › Spatiotemporal Contrastive Video Representation Learning

Spatiotemporal Contrastive Video Representation Learning

9 Aug 2020CVPR 2021 1arXiv:2008.03800archive 2025-07-28

Rui Qian, Tianjian Meng, Boqing Gong, Ming-Hsuan Yang, Huisheng Wang, Serge Belongie, Yin Cui

We present a self-supervised Contrastive Video Representation Learning (CVRL) method to learn spatiotemporal visual representations from unlabeled videos. Our representations are learned using a contrastive loss, where two augmented clips from the same short video are pulled together in the embedding space, while clips from different videos are pushed away. We study what makes for good data augmentations for video self-supervised learning and find that both spatial and temporal information are crucial. We carefully design data augmentations involving spatial and temporal cues. Concretely, we propose a temporally consistent spatial augmentation method to impose strong spatial augmentations on each frame of the video while maintaining the temporal consistency across frames. We also propose a sampling-based temporal augmentation method to avoid overly enforcing invariance on clips that are distant in time. On Kinetics-600, a linear classifier trained on the representations learned by CVRL achieves 70.4% top-1 accuracy with a 3D-ResNet-50 (R3D-50) backbone, outperforming ImageNet supervised pre-training by 15.7% and SimCLR unsupervised pre-training by 18.8% using the same inflated R3D-50. The performance of CVRL can be further improved to 72.9% with a larger R3D-152 (2x filters) backbone, significantly closing the gap between unsupervised and supervised video representation learning. Our code and models will be available at https://github.com/tensorflow/models/tree/master/official/.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

applecrumble123/CVLR_pytorch mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionContrastive LearningData AugmentationRepresentation LearningSelf-Supervised Action RecognitionSelf-Supervised Action Recognition LinearSelf-Supervised LearningUnsupervised Pre-training

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Self-Supervised Action Recognition HMDB51 CVRL (R3D-152 2x; K600) Frozen false #7 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 CVRL (R3D-152 2x; K600) Pre-Training Dataset Kinetics600 #7 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 CVRL (R3D-152 2x; K600) Top-1 Accuracy 69.9 #7 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 CVRL (R3D-50; K600) Frozen false #10 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 CVRL (R3D-50; K600) Pre-Training Dataset Kinetics600 #10 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 CVRL (R3D-50; K600) Top-1 Accuracy 68.0 #10 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 CVRL (R3D-50; K400) Frozen false #12 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 CVRL (R3D-50; K400) Pre-Training Dataset Kinetics400 #12 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 CVRL (R3D-50; K400) Top-1 Accuracy 66.7 #12 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 (finetuned) CVRL (R3D-152 2x; K600) Pretraining Dataset K600 #3 of 14 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 (finetuned) CVRL (R3D-152 2x; K600) Top-1 Accuracy 69.9 #3 of 14 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 (finetuned) CVRL (R3D-50; K600) Pretraining Dataset K600 #5 of 14 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 (finetuned) CVRL (R3D-50; K600) Top-1 Accuracy 68.0 #5 of 14 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 (finetuned) CVRL (R3D-50; K400) Pretraining Dataset K400 #7 of 14 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 (finetuned) CVRL (R3D-50; K400) Top-1 Accuracy 66.7 #7 of 14 Archive leaderboard report
Self-Supervised Action Recognition Kinetics-400 CVRL (R3D-152 2x; K600 pretrain) Top-1 accuracy % 71.6 #2 of 4 Archive leaderboard report
Self-Supervised Action Recognition Kinetics-400 CVRL (R3D-101) Top-1 accuracy % 67.6 #3 of 4 Archive leaderboard report
Self-Supervised Action Recognition Kinetics-400 CVRL (R3D-50) Top-1 accuracy % 66.1 #4 of 4 Archive leaderboard report
Self-Supervised Action Recognition Kinetics-600 CVRL (R3D-152 2x) Top-1 Accuracy 72.9 #1 of 5 Archive leaderboard report
Self-Supervised Action Recognition Kinetics-600 CVRL (R3D-101) Top-1 Accuracy 71.6 #2 of 5 Archive leaderboard report
Self-Supervised Action Recognition Kinetics-600 CVRL (R3D-50) Top-1 Accuracy 70.4 #4 of 5 Archive leaderboard report
Self-Supervised Action Recognition UCF101 CVRL (R3D-152 2x; K600) 3-fold Accuracy 93.9 #10 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 CVRL (R3D-152 2x; K600) Frozen false #10 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 CVRL (R3D-152 2x; K600) Pre-Training Dataset Kinetics600 #10 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 CVRL (R3D-50; K600) 3-fold Accuracy 93.4 #12 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 CVRL (R3D-50; K600) Frozen false #12 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 CVRL (R3D-50; K600) Pre-Training Dataset Kinetics600 #12 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 CVRL (R3D-50; K400) 3-fold Accuracy 92.2 #17 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 CVRL (R3D-50; K400) Frozen false #17 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 CVRL (R3D-50; K400) Pre-Training Dataset Kinetics400 #17 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 (finetuned) CVRL (R3D-152 2x; K600) 3-fold Accuracy 93.9 #3 of 14 Archive leaderboard report
Self-Supervised Action Recognition UCF101 (finetuned) CVRL (R3D-152 2x; K600) Pretrain K600 #3 of 14 Archive leaderboard report
Self-Supervised Action Recognition UCF101 (finetuned) CVRL (R3D-50; K600) 3-fold Accuracy 93.4 #5 of 14 Archive leaderboard report
Self-Supervised Action Recognition UCF101 (finetuned) CVRL (R3D-50; K600) Pretrain K600 #5 of 14 Archive leaderboard report
Self-Supervised Action Recognition UCF101 (finetuned) CVRL (R3D-50; K400) 3-fold Accuracy 92.2 #6 of 14 Archive leaderboard report
Self-Supervised Action Recognition UCF101 (finetuned) CVRL (R3D-50; K400) Pretrain K400 #6 of 14 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Introduced by this paper: CVRL

1x1 Convolution3D ConvolutionAverage PoolingBatch NormalizationBottleneck Residual BlockCVRLColorJitterConvolutionDense ConnectionsFeedforward NetworkGlobal Average PoolingKaiming InitializationMax PoolingNT-XentRandom Gaussian BlurRandom Resized CropReLUResidual BlockResidual ConnectionSimCLRTemporally Consistent Spatial Augmentation

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections