Browse State-of-the-Art › 3D Action Recognition
3D Action Recognition
38 papers with code · 3 benchmarks · 15 datasets archive 2025-07-28
Image: Rahmani et al
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Assembly101 (7 rows) | HandFormer-B/21 | On the Utility of 3D Hand Poses for Action Recognition | code | Syntology ran 8 of 13 samples · 5 unverified | Compare |
| NTU RGB+D (5 rows) | Kinet | No Pain, Big Gain: Classify Dynamic Point Cloud Sequences with... | code | — | Compare |
| 100 sleep nights of 8 caregivers (1 row) | htf | A System for Real-Time Interactive Analysis of Deep Learning Training | code | Syntology ran 0 of 17 samples · 17 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
15 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
6 subtasks in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 38 papers with code (91 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
30 Oct 2017 24 repositories listed Syntology ran 4 of 7 samples · 3 unverified · 3 pointer-only (licence)Over the last decade, Convolutional Neural Network (CNN) models have been highly successful in solving complex vision problems.
-
20 Nov 2018 13 repositories listed Syntology ran 6 of 16 samples · 10 unverified · 4 pointer-only (licence)The explosive growth in video streaming gives rise to challenges on performing video understanding at high accuracy and low computation cost.
-
19 Jun 2019 6 repositories listed Syntology ran 4 of 11 samples · 7 unverified · 4 pointer-only (licence)In this work we aim to learn object representations that are useful for control and reinforcement learning (RL).
-
28 Apr 2021 4 repositories listedIn this work, we propose PoseC3D, a new approach to skeleton-based action recognition, which relies on a 3D heatmap stack instead of a graph sequence as the base representation of human skeletons.
-
20 May 2018 4 repositories listedIn addition, the second-order information (the lengths and directions of bones) of the skeleton data, which is naturally more informative and discriminative for action recognition, is rarely investigated in existing…
-
31 Mar 2020 3 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedSpatial-temporal graphs have been widely used by skeleton-based action recognition algorithms to model human action dynamics.
-
30 May 2020 2 repositories listedOne primary technical challenge in photoacoustic microscopy (PAM) is the necessary compromise between spatial resolution and imaging speed.
-
11 Apr 2016 2 repositories listedRecent approaches in depth-based human activity analysis achieved outstanding performance and proved the effectiveness of 3D representation for classification of action classes.
-
9 Oct 2024 1 repository listed Syntology ran 2 of 2 samples · 0 unverifiedTo this end, we introduce a Convex Hull Adaptive Shift based multi-Entity action recognition method (CHASE), which mitigates inter-entity distribution gaps and unbiases subsequent backbones.
-
26 Sep 2024 1 repository listedIn self-supervised skeleton-based action recognition, the mask reconstruction paradigm is gaining interest in enhancing model refinement and robustness through effective masking.
-
15 Jul 2024 1 repository listedSelf-supervised pretraining methods with masked prediction demonstrate remarkable within-dataset performance in skeleton-based action recognition.
-
14 Mar 2024 1 repository listed Syntology ran 8 of 13 samples · 5 unverified3D hand pose is an underexplored modality for action recognition.
-
14 Aug 2023 1 repository listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)To be specific, the proposed MAMP takes as input the masked spatio-temporal skeleton sequence and predicts the corresponding temporal motion of the masked human joints.
-
14 Jul 2023 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedTo address these problems, we propose an Interactive Spatiotemporal Token Attention Network (ISTA-Net), which simultaneously model spatial, temporal, and interactive relations.
-
26 Aug 2022 1 repository listedIn this work, we formulate the cross-modal interaction as a bidirectional knowledge distillation problem.
-
20 Jul 2022 1 repository listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)Furthermore, to leverage the complementarity of domain-shared features and target-specific features, we propose a novel collaborative clustering strategy to enforce pair-wise relationship consistency between the two…
-
27 Jun 2022 1 repository listed Syntology ran 19 of 26 samples · 7 unverifiedTo solve this problem, we present a multi-scale spatial graph convolution (MS-GC) module and a multi-scale temporal graph convolution (MT-GC) module to enrich the receptive field of the model in spatial and temporal…
-
27 May 2022 1 repository listedThen, a spatial convolution is employed to capture the local structure of points in the 3D space, and a temporal convolution is used to model the dynamics of the spatial regions along the time dimension.
-
28 Mar 2022 1 repository listedAssembly101 is a new procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 "take-apart" toy vehicles.
-
21 Mar 2022 1 repository listedScene flow is a powerful tool for capturing the motion field of 3D point clouds.
-
28 Feb 2022 1 repository listedTo that end, we propose to learn exercise-oriented image and video representations from unlabeled samples such that a small dataset annotated by experts suffices for supervised error detection.
-
14 Dec 2021 1 repository listedSpecifically, a spatial operation is employed to capture the local structure of each spatial region in a tube and a temporal operation is used to model the dynamics of the spatial regions along the tube.
-
1 Dec 2021 1 repository listedFurthermore, unlike existing volumetric MVS techniques, our 3D CNN operates on a feature-augmented point cloud, allowing for effective aggregation of multi-view information and flexible iterative refinement of depth…
-
16 Nov 2021 1 repository listedInstead of capturing spatio-temporal local structures, SequentialPointNet encodes the temporal evolution of static appearances to recognize human actions.
-
19 Jun 2021 1 repository listedTo capture the dynamics in point cloud videos, point tracking is usually employed.
-
17 Jun 2021 1 repository listedTo address this, we present BABEL, a large dataset with language labels describing the actions being performed in mocap sequences.
-
1 Aug 2020 1 repository listedWe propose novel approaches for simultaneously identifying important weights of a convolutional neural network (ConvNet) and providing more attention to the important weights during training.
-
Roweisposes, Including Eigenposes, Supervised Eigenposes, and Fisherposes, for 3D Action Recognition28 Jun 2020 1 repository listedAlthough various methods have been proposed for 3D action recognition, some of which are basic and some use deep learning, the need of basic methods based on generalized eigenvalue problem is sensed for action…
-
12 May 2020 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Each available 3DV voxel intrinsically involves 3D spatial and motion feature jointly.
-
5 Jan 2020 1 repository listed Syntology ran 0 of 17 samples · 17 unverifiedTo achieve this, we model various exploratory inspection and diagnostic tasks for deep learning training processes as specifications for streams using a map-reduce paradigm with which many data scientists are already…
Syntology lines on 12 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections