Papers › MaskCLR: Attention-Guided Contrastive Learning for Robust Action Representation Learning

MaskCLR: Attention-Guided Contrastive Learning for Robust Action Representation Learning

1 Jan 2024CVPR 2024 1archive 2025-07-28

Mohamed Abdelfattah, Mariam Hassan, Alexandre Alahi

Current transformer-based skeletal action recognition models tend to focus on a limited set of joints and low-level motion patterns to predict action classes. This results in significant performance degradation under small skeleton perturbations or changing the pose estimator between training and testing. In this work we introduce MaskCLR a new Masked Contrastive Learning approach for Robust skeletal action recognition. We propose an Attention-Guided Probabilistic Masking strategy to occlude the most important joints and encourage the model to explore a larger set of discriminative joints. Furthermore we propose a Multi-Level Contrastive Learning paradigm to enforce the representations of standard and occluded skeletons to be class-discriminative i.e. more compact within each class and more dispersed across different classes. Our approach helps the model capture the high-level action semantics instead of low-level joint variations and can be conveniently incorporated into transformer-based models. Without loss of generality we combine MaskCLR with three transformer backbones: the vanilla transformer DSTFormer and STTFormer. Extensive experiments on NTU60 NTU120 and Kinetics400 show that MaskCLR consistently outperforms previous state-of-the-art methods on standard and perturbed skeletons from different pose estimators showing improved accuracy generalization and robustness. Project website: https://maskclr.github.io.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionContrastive LearningRepresentation LearningSkeleton Based Action Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Skeleton Based Action Recognition NTU RGB+D MaskCLR Accuracy (CS) 93.9 #4 of 135 Archive leaderboard report
Skeleton Based Action Recognition NTU RGB+D MaskCLR Accuracy (CV) 97.3 #4 of 135 Archive leaderboard report
Skeleton Based Action Recognition NTU RGB+D 120 MaskCLR Accuracy (Cross-Setup) 89.5 #33 of 83 Archive leaderboard report
Skeleton Based Action Recognition NTU RGB+D 120 MaskCLR Accuracy (Cross-Subject) 87.4 #33 of 83 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Contrastive LearningFocusSET

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections