Papers › Interpretable 3D Human Action Analysis with Temporal Convolutional Networks

Interpretable 3D Human Action Analysis with Temporal Convolutional Networks

14 Apr 2017arXiv:1704.04516archive 2025-07-28

Tae Soo Kim, Austin Reiter

The discriminative power of modern deep learning models for 3D human action recognition is growing ever so potent. In conjunction with the recent resurgence of 3D human action representation with 3D skeletons, the quality and the pace of recent progress have been significant. However, the inner workings of state-of-the-art learning based methods in 3D human action recognition still remain mostly black-box. In this work, we propose to use a new class of models known as Temporal Convolutional Neural Networks (TCN) for 3D human action recognition. Compared to popular LSTM-based Recurrent Neural Network models, given interpretable input such as 3D skeletons, TCN provides us a way to explicitly learn readily interpretable spatio-temporal representations for 3D human action recognition. We provide our strategy in re-designing the TCN with interpretability in mind and how such characteristics of the model is leveraged to construct a powerful 3D activity recognition method. Through this work, we wish to take a step towards a spatio-temporal model that is easier to understand, explain and interpret. The resulting model, Res-TCN, achieves state-of-the-art results on the largest 3D human action recognition dataset, NTU-RGBD.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

TaeSoo-Kim/TCNActionRecognition officialmentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

3D Action RecognitionAction AnalysisAction RecognitionActivity RecognitionMultimodal Activity RecognitionSkeleton Based Action RecognitionTemporal Action Localization

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Multimodal Activity Recognition EV-Action TCN (Skeleton Kinect) Accuracy 80.1 #1 of 9 Archive leaderboard report
Multimodal Activity Recognition EV-Action TCN (Skeleton Vicon) Accuracy 64.1 #6 of 9 Archive leaderboard report
Skeleton Based Action Recognition NTU RGB+D TCN Accuracy (CS) 74.3 #123 of 135 Archive leaderboard report
Skeleton Based Action Recognition NTU RGB+D TCN Accuracy (CV) 83.1 #123 of 135 Archive leaderboard report
Skeleton Based Action Recognition Varying-view RGB-D Action-Skeleton Res-TCN Accuracy (AV I) 48% #3 of 7 Archive leaderboard report
Skeleton Based Action Recognition Varying-view RGB-D Action-Skeleton Res-TCN Accuracy (AV II) 68% #3 of 7 Archive leaderboard report
Skeleton Based Action Recognition Varying-view RGB-D Action-Skeleton Res-TCN Accuracy (CS) 63% #3 of 7 Archive leaderboard report
Skeleton Based Action Recognition Varying-view RGB-D Action-Skeleton Res-TCN Accuracy (CV I) 14% #3 of 7 Archive leaderboard report
Skeleton Based Action Recognition Varying-view RGB-D Action-Skeleton Res-TCN Accuracy (CV II) 48% #3 of 7 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections