Papers › SkeleTR: Towards Skeleton-based Action Recognition in the Wild

SkeleTR: Towards Skeleton-based Action Recognition in the Wild

1 Jan 2023ICCV 2023 1archive 2025-07-28

Haodong Duan, Mingze Xu, Bing Shuai, Davide Modolo, Zhuowen Tu, Joseph Tighe, Alessandro Bergamo

We present SkeleTR, a new framework for skeleton-based action recognition. In contrast to prior work, which focuses mainly on controlled environments, we target in-the-wild scenarios that typically involve a variable number of people and various forms of interaction between people. SkeleTR works with a two-stage paradigm. It first models the intra-person skeleton dynamics for each skeleton sequence with graph convolutions, and then uses stacked Transformer encoders to capture person interactions that are important for action recognition in the wild. To mitigate the negative impact of inaccurate skeleton associations, SkeleTR takes relative short skeleton sequences as input and increases the number of sequences. As a unified solution, SkeleTR can be directly applied to multiple skeleton-based action tasks, including video-level action classification, instance-level action detection, and group-level activity recognition. It also enables transfer learning and joint training across different action tasks and datasets, which results in performance improvement. When evaluated on various skeleton-based action recognition benchmarks, SkeleTR achieves the state-of-the-art performance.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action ClassificationAction DetectionAction RecognitionActivity RecognitionHuman Interaction RecognitionSkeleton Based Action RecognitionTransfer Learning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Human Interaction Recognition NTU RGB+D SkeleTR Accuracy (Cross-Subject) 94.9 #3 of 5 Archive leaderboard report
Human Interaction Recognition NTU RGB+D SkeleTR Accuracy (Cross-View) 97.7 #3 of 5 Archive leaderboard report
Human Interaction Recognition NTU RGB+D 120 SkeleTR Accuracy (Cross-Setup) 88.3 #4 of 6 Archive leaderboard report
Human Interaction Recognition NTU RGB+D 120 SkeleTR Accuracy (Cross-Subject) 87.8 #4 of 6 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Absolute Position EncodingsAdamAttentionBPEDense ConnectionsDropoutLabel SmoothingLayer NormalizationLinear LayerMulti-Head AttentionPosition-Wise Feed-Forward LayerResidual ConnectionSoftmaxTransformer

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections