Browse State-of-the-Art › Self-Supervised Action Recognition
Self-Supervised Action Recognition
35 papers with code · 6 benchmarks · 5 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
6 leaderboard tables shown for this task, 6 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| UCF101 (53 rows) | VideoMAE V2-g | VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking | code | Syntology ran 2 of 6 samples · 4 unverified | Compare |
| HMDB51 (48 rows) | MVD (ViT-B) | Masked Video Distillation: Rethinking Masked Feature Modeling for... | code | — | Compare |
| HMDB51 (finetuned) (14 rows) | BraVe:V-FA (TSM-50x2) | Broaden Your Views for Self-Supervised Video Learning | code | Syntology ran 0 of 8 samples · 8 unverified | Compare |
| UCF101 (finetuned) (14 rows) | BraVe:V-FA (TSM-50x2) | Broaden Your Views for Self-Supervised Video Learning | code | Syntology ran 0 of 8 samples · 8 unverified | Compare |
| Kinetics-600 (5 rows) | CVRL (R3D-152 2x) | Spatiotemporal Contrastive Video Representation Learning | code | — | Compare |
| Kinetics-400 (4 rows) | XKD (ViT-B/112/16) | XKD: Cross-modal Knowledge Distillation with Domain Alignment for... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
5 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 35 papers with code (51 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
23 Mar 2022 9 repositories listed Syntology ran 9 of 13 samples · 4 unverified · 12 pointer-only (licence)Pre-training video transformers on extra large-scale datasets is generally required to achieve premier performance on relatively small datasets.
-
13 Jun 2019 8 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 1 pointer-only (licence)We analyze key properties of the approach that make it work, finding that the contrastive loss outperforms a popular alternative based on cross-view prediction, and that the more views we learn from, the better the…
-
8 Dec 2022 4 repositories listedFor the choice of teacher models, we observe that students taught by video teachers perform better on temporally-heavy video tasks, while image teachers transfer stronger spatial representations for spatially-heavy…
-
9 Aug 2020 4 repositories listedOur representations are learned using a contrastive loss, where two augmented clips from the same short video are pulled together in the embedding space, while clips from different videos are pushed away.
-
15 Jul 2023 2 repositories listedInspired by SkeletonBYOL, this paper further presents a Cross-Model and Cross-Stream (CMCS) framework.
-
21 Dec 2022 2 repositories listedA good data representation should contain relations between the instances, or semantic similarity and dissimilarity, that contrastive learning harms by considering all negatives as noise.
-
19 Nov 2022 2 repositories listed Syntology ran 3 of 6 samples · 3 unverified · 6 pointer-only (licence)Masked Video Autoencoder (MVA) approaches have demonstrated their potential by significantly outperforming previous video representation learning methods.
-
12 Oct 2022 2 repositories listedThe latest attempts seek to learn a representation model by predicting the appearance contents in the masked regions.
-
29 Apr 2021 2 repositories listedWe present a large-scale study on unsupervised spatiotemporal representation learning from videos.
-
6 Aug 2020 2 repositories listed Syntology ran 4 of 4 samples · 0 unverified · 4 pointer-only (licence)With the proposed Inter-Intra Contrastive (IIC) framework, we can train spatio-temporal convolutional networks to learn video representations.
-
1 May 2023 1 repository listed Syntology ran 6 of 12 samples · 6 unverified · 12 pointer-only (licence)This paper proposes an attention-based contrastive learning framework for skeleton representation learning, called SkeAttnCLR, which integrates local similarity and global features for skeleton-based action…
-
29 Mar 2023 1 repository listed Syntology ran 2 of 6 samples · 4 unverifiedFinally, we successfully train a video ViT model with a billion parameters, which achieves a new state-of-the-art performance on the datasets of Kinetics (90.
-
17 Feb 2023 1 repository listed Syntology ran 3 of 5 samples · 2 unverifiedSpecifically, we construct a negative-sample-free triplet steam structure that is composed of an anchor stream without any masking, a spatial masking stream with Central Spatial Masking (CSM), and a temporal masking…
-
25 Nov 2022 1 repository listedFirst, masked data reconstruction is performed to learn modality-specific representations from audio and visual streams.
-
25 Jun 2022 1 repository listedOne of the key reasons for this is that sampling pairs of similar video clips, a required step for many self-supervised contrastive learning methods, is currently done conservatively to avoid false positives.
-
7 Dec 2021 1 repository listedIn this paper, to make better use of the movement patterns introduced by extreme augmentations, a Contrastive Learning framework utilizing Abundant Information Mining for self-supervised action Representation (AimCLR)…
-
9 Nov 2021 1 repository listedWe present CrissCross, a self-supervised framework for learning audio-visual representations.
-
19 Jun 2021 1 repository listed Syntology ran 0 of 4 samples · 4 unverifiedWe introduce a framework for learning from unlabeled video what is predictable in the future.
-
18 Jun 2021 1 repository listedInstance-level contrastive learning techniques, which rely on data augmentation and a contrastive loss function, have found great success in the domain of visual representation learning.
-
30 Mar 2021 1 repository listed Syntology ran 0 of 8 samples · 8 unverifiedMost successful self-supervised learning methods are trained to align the representations of two independent views from the data.
-
20 Jan 2021 1 repository listed Syntology ran 0 of 1 samples · 1 unverifiedHowever, prior work on contrastive learning for video data has not explored the effect of explicitly encouraging the features to be distinct across the temporal dimension.
-
29 Oct 2020 1 repository listedIt is convenient to treat PCL as a standard training strategy and apply it to many other works in self-supervised video feature learning.
-
27 Oct 2020 1 repository listedWe study unsupervised video representation learning that seeks to learn both motion and appearance features from unlabeled video only, which can be reused for downstream tasks such as action recognition.
-
19 Oct 2020 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedThe objective of this paper is visual-only self-supervised video representation learning.
-
29 Jun 2020 1 repository listedIn particular, we explore how best to combine the modalities, such that fine-grained representations of the visual and audio modalities can be maintained, whilst also integrating text into a common embedding.
-
1 Jun 2020 1 repository listedThe generative perception model acts as a feature decoder to focus on comprehending high temporal resolution and short-term representation by introducing a motion-attention mechanism.
-
27 Apr 2020 1 repository listedOur method uses contrastive learning for cross-modal discrimination of video from audio and vice-versa.
-
13 Apr 2020 1 repository listedWe demonstrate how those learned features can boost the performance of self-supervised action recognition, and can be used for video retrieval.
-
21 Mar 2020 1 repository listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)The proposed method exploits inherent structure of unlabeled video data to explicitly enforce temporal coherency in the embedding space, rather than indirectly learning it through ranking or predictive proxy tasks.
-
5 Mar 2020 1 repository listedWe propose a self-supervised visual learning method by predicting the variable playback speeds of a video.
Syntology lines on 12 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections