Papers › Skeleton-Based Action Recognition Using Spatio-Temporal LSTM Network with Trust Gates

Skeleton-Based Action Recognition Using Spatio-Temporal LSTM Network with Trust Gates

26 Jun 2017arXiv:1706.08276archive 2025-07-28

Jun Liu, Amir Shahroudy, Dong Xu, Alex C. Kot, Gang Wang

Skeleton-based human action recognition has attracted a lot of research attention during the past few years. Recent works attempted to utilize recurrent neural networks to model the temporal dependencies between the 3D positional configurations of human body joints for better analysis of human activities in the skeletal data. The proposed work extends this idea to spatial domain as well as temporal domain to better analyze the hidden sources of action-related information within the human skeleton sequences in both of these domains simultaneously. Based on the pictorial structure of Kinect's skeletal data, an effective tree-structure based traversal framework is also proposed. In order to deal with the noise in the skeletal data, a new gating mechanism within LSTM module is introduced, with which the network can learn the reliability of the sequential data and accordingly adjust the effect of the input data on the updating procedure of the long-term context representation stored in the unit's memory cell. Moreover, we introduce a novel multi-modal feature fusion strategy within the LSTM unit in this paper. The comprehensive experimental results on seven challenging benchmark datasets for human action recognition demonstrate the effectiveness of the proposed method.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionOne-Shot 3D Action RecognitionSkeleton Based Action RecognitionTemporal Action Localization

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
One-Shot 3D Action Recognition NTU RGB+D 120 Average Pooling Accuracy 42.9% #8 of 10 Archive leaderboard report
Skeleton Based Action Recognition NTU RGB+D 120 Internal Feature Fusion Accuracy (Cross-Setup) 60.9% #79 of 83 Archive leaderboard report
Skeleton Based Action Recognition NTU RGB+D 120 Internal Feature Fusion Accuracy (Cross-Subject) 58.2% #79 of 83 Archive leaderboard report
Skeleton Based Action Recognition SYSU 3D ST-LSTM (Tree) Accuracy 73.4% #9 of 9 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

LSTMSigmoid ActivationTanh Activation

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections