Papers › Tensor Representations for Action Recognition

Tensor Representations for Action Recognition

28 Dec 2020arXiv:2012.14371archive 2025-07-28

Piotr Koniusz, Lei Wang, Anoop Cherian

Human actions in video sequences are characterized by the complex interplay between spatial features and their temporal dynamics. In this paper, we propose novel tensor representations for compactly capturing such higher-order relationships between visual features for the task of action recognition. We propose two tensor-based feature representations, viz. (i) sequence compatibility kernel (SCK) and (ii) dynamics compatibility kernel (DCK). SCK builds on the spatio-temporal correlations between features, whereas DCK explicitly models the action dynamics of a sequence. We also explore generalization of SCK, coined SCK(+), that operates on subsequences to capture the local-global interplay of correlations, which can incorporate multi-modal inputs e.g., skeleton 3D body-joints and per-frame classifier scores obtained from deep learning models trained on videos. We introduce linearization of these kernels that lead to compact and fast descriptors. We provide experiments on (i) 3D skeleton action sequences, (ii) fine-grained video sequences, and (iii) standard non-fine-grained videos. As our final representations are tensors that capture higher-order relationships of features, they relate to co-occurrences for robust fine-grained recognition. We use higher-order tensors and so-called Eigenvalue Power Normalization (EPN) which have been long speculated to perform spectral detection of higher-order occurrences, thus detecting fine-grained relationships of features rather than merely count features in action sequences. We prove that a tensor of order r, built from Z* dimensional features, coupled with EPN indeed detects if at least one higher-order occurrence is `projected' into one of its binom(Z*,r) subspaces of dim. r represented by the tensor, thus forming a Tensor Power Normalization metric endowed with binom(Z*,r) such `detectors'.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionAction Recognition In VideosSkeleton Based Action Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Recognition HMDB-51 SCK⊕(I3D)+IDT Average accuracy of 3 splits 86.11 #5 of 77 Archive leaderboard report
Skeleton Based Action Recognition Florence 3D SCK⊕+DCK⊕ Accuracy 97.45 #3 of 7 Archive leaderboard report
Skeleton Based Action Recognition Florence 3D SCK+DCK Accuracy 95.23 #5 of 7 Archive leaderboard report
Skeleton Based Action Recognition NTU RGB+D SCK⊕ Accuracy (CS) 91.56 #38 of 135 Archive leaderboard report
Skeleton Based Action Recognition NTU RGB+D SCK⊕ Accuracy (CV) 94.75 #38 of 135 Archive leaderboard report
Skeleton Based Action Recognition UT-Kinect SCK⊕+DCK⊕ Accuracy 99.2 #2 of 7 Archive leaderboard report
Skeleton Based Action Recognition UT-Kinect SCK+DCK Accuracy 98.2 #5 of 7 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections