Papers › Self-supervised Video Representation Learning with Cross-Stream Prototypical Contrasting

Self-supervised Video Representation Learning with Cross-Stream Prototypical Contrasting

18 Jun 2021arXiv:2106.10137archive 2025-07-28

Martine Toering, Ioannis Gatopoulos, Maarten Stol, Vincent Tao Hu

Instance-level contrastive learning techniques, which rely on data augmentation and a contrastive loss function, have found great success in the domain of visual representation learning. They are not suitable for exploiting the rich dynamical structure of video however, as operations are done on many augmented instances. In this paper we propose "Video Cross-Stream Prototypical Contrasting", a novel method which predicts consistent prototype assignments from both RGB and optical flow views, operating on sets of samples. Specifically, we alternate the optimization process; while optimizing one of the streams, all views are mapped to one set of stream prototype vectors. Each of the assignments is predicted with all views except the one matching the prediction, pushing representations closer to their assigned prototypes. As a result, more efficient video embeddings with ingrained motion information are learned, without the explicit need for optical flow computation during inference. We obtain state-of-the-art results on nearest-neighbour video retrieval and action recognition, outperforming previous best by +3.2% on UCF101 using the S3D backbone (90.5% Top-1 acc), and by +7.2% on UCF101 and +15.1% on HMDB51 using the R(2+1)D backbone.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

martinetoering/ViCC officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionAction Recognition In VideosContrastive LearningData AugmentationOptical Flow EstimationRepresentation LearningRetrievalSelf-Supervised Action RecognitionSelf-Supervised LearningSelf-supervised Video RetrievalVideo ClassificationVideo RecognitionVideo Retrieval

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Self-Supervised Action Recognition HMDB51 ViCC (S3D; R+F) Frozen false #23 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 ViCC (S3D; R+F) Pre-Training Dataset UCF101 #23 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 ViCC (S3D; R+F) Top-1 Accuracy 62.2 #23 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 ViCC (R2+1D; R+F) Frozen false #24 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 ViCC (R2+1D; R+F) Pre-Training Dataset UCF101 #24 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 ViCC (R2+1D; R+F) Top-1 Accuracy 61.5 #24 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 ViCC (R2+1D; RGB) Frozen false #33 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 ViCC (R2+1D; RGB) Pre-Training Dataset UCF101 #33 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 ViCC (R2+1D; RGB) Top-1 Accuracy 52.4 #33 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 ViCC (S3D; RGB) Frozen true #36 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 ViCC (S3D; RGB) Pre-Training Dataset UCF101 #36 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 ViCC (S3D; RGB) Top-1 Accuracy 38.5 #36 of 48 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 (finetuned) ViCC (S3D; R+F) Pretraining Dataset UCF101 #9 of 14 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 (finetuned) ViCC (S3D; R+F) Top-1 Accuracy 62.2 #9 of 14 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 (finetuned) ViCC (R2+1D; RGB) Pretraining Dataset UCF101 #13 of 14 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 (finetuned) ViCC (R2+1D; RGB) Top-1 Accuracy 52.4 #13 of 14 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 (finetuned) ViCC (S3D; RGB)) Pretraining Dataset UCF101 #14 of 14 Archive leaderboard report
Self-Supervised Action Recognition HMDB51 (finetuned) ViCC (S3D; RGB)) Top-1 Accuracy 47.9 #14 of 14 Archive leaderboard report
Self-Supervised Action Recognition UCF101 ViCC (S3D; R+F) 3-fold Accuracy 90.5 #22 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 ViCC (S3D; R+F) Frozen false #22 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 ViCC (S3D; R+F) Pre-Training Dataset UCF101 #22 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 ViCC (S3D; RGB) 3-fold Accuracy 88.8 #23 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 ViCC (S3D; RGB) Frozen false #23 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 ViCC (S3D; RGB) Pre-Training Dataset UCF101 #23 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 ViCC (R2+1D; R+F) 3-fold Accuracy 88.8 #24 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 ViCC (R2+1D; R+F) Frozen false #24 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 ViCC (R2+1D; R+F) Pre-Training Dataset UCF101 #24 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 ViCC (R2+1D; RGB) 3-fold Accuracy 82.8 #30 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 ViCC (R2+1D; RGB) Frozen false #30 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 ViCC (R2+1D; RGB) Pre-Training Dataset UCF101 #30 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 ViCC (S3D; RGB) 3-fold Accuracy 72.2 #36 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 ViCC (S3D; RGB) Frozen true #36 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 ViCC (S3D; RGB) Pre-Training Dataset UCF101 #36 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 (finetuned) ViCC (S3D; R+F) 3-fold Accuracy 90.5 #9 of 14 Archive leaderboard report
Self-Supervised Action Recognition UCF101 (finetuned) ViCC (S3D; R+F) Pretrain UCF101 #9 of 14 Archive leaderboard report
Self-Supervised Action Recognition UCF101 (finetuned) ViCC (R2+1D; R+F) 3-fold Accuracy 88.8 #11 of 14 Archive leaderboard report
Self-Supervised Action Recognition UCF101 (finetuned) ViCC (R2+1D; R+F) Pretrain UCF101 #11 of 14 Archive leaderboard report
Self-Supervised Action Recognition UCF101 (finetuned) ViCC (S3D; RGB) 3-fold Accuracy 84.3 #13 of 14 Archive leaderboard report
Self-Supervised Action Recognition UCF101 (finetuned) ViCC (S3D; RGB) Pretrain UCF101 #13 of 14 Archive leaderboard report
Self-Supervised Action Recognition UCF101 (finetuned) ViCC (R2+1D; RGB) 3-fold Accuracy 82.8 #14 of 14 Archive leaderboard report
Self-Supervised Action Recognition UCF101 (finetuned) ViCC (R2+1D; RGB) Pretrain UCF101 #14 of 14 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

(2+1)D Convolution3D ConvolutionAverage PoolingBatch NormalizationContrastive LearningDense ConnectionsGlobal Average PoolingR(2+1)DReLUResidual Connection

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections