Papers › MSA-GCN: Exploiting Multi-Scale Temporal Dynamics With Adaptive Graph Convolution for...

MSA-GCN: Exploiting Multi-Scale Temporal Dynamics With Adaptive Graph Convolution for Skeleton-Based Action Recognition

19 Dec 2024IEEE Access 2024 12archive 2025-07-28

Kowovi Comivi Alowonou, Ji-Hyeong Han

Graph convolutional networks (GCNs) have been widely used and have achieved remarkable results in skeleton-based action recognition. We note that existing GCN-based approaches rely on local context information of the skeleton joints to construct adaptive graphs for feature aggregation, limiting their ability to understand actions that involve coordinated movements across various parts of the body. An adaptive graph built upon the global context information of the joints can help move beyond this limitation. Therefore, in this paper, we propose a novel approach to skeleton-based action recognition named Multi-stage Adaptive Graph Convolution Network (MSA-GCN). It consists of two modules: Multi-stage Adaptive Graph Convolution (MSA-GC) and Temporal Multi-Scale Transformer (TMST). These two modules work together to capture complex spatial and temporal patterns within skeleton data effectively. Specifically, MSA-GC explores both local and global context information of the joints across all sequences to construct the adaptive graph and facilitates the understanding of complex and nuanced relationships between joints. On the other hand, the TMST module integrates a Gated Multi-stage Temporal Convolution (GMSTC) with a Temporal Multi-Head Self-Attention (TMHSA) to capture global temporal features and accommodate both long-term and short-term dependencies within action sequences. Through extensive experiments on multiple benchmark datasets, including NTU RGB+D 60, NTU RGB+D 120, and Northwestern-UCLA, MSA-GCN achieves state-of-the-art performance and verifies its effectiveness in skeleton-based action recognition.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionSkeleton Based Action Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Skeleton Based Action Recognition NTU RGB+D MSA-GCN Accuracy (CS) 93.6 #8 of 135 Archive leaderboard report
Skeleton Based Action Recognition NTU RGB+D MSA-GCN Accuracy (CV) 97.4 #8 of 135 Archive leaderboard report
Skeleton Based Action Recognition NTU RGB+D MSA-GCN Ensembled Modalities 6 #8 of 135 Archive leaderboard report
Skeleton Based Action Recognition NTU RGB+D 120 MSA-GCN Accuracy (Cross-Setup) 92.2 #4 of 83 Archive leaderboard report
Skeleton Based Action Recognition NTU RGB+D 120 MSA-GCN Accuracy (Cross-Subject) 90.6 #4 of 83 Archive leaderboard report
Skeleton Based Action Recognition NTU RGB+D 120 MSA-GCN Ensembled Modalities 6 #4 of 83 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEConvolutionDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections