Papers › Deep set conditioned latent representations for action recognition

Deep set conditioned latent representations for action recognition

21 Dec 2022International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications 2022 2arXiv:2212.11030archive 2025-07-28

Akash Singh, Tom De Schepper, Kevin Mets, Peter Hellinckx, Jose Oramas, Steven Latre

In recent years multi-label, multi-class video action recognition has gained significant popularity. While reasoning over temporally connected atomic actions is mundane for intelligent species, standard artificial neural networks (ANN) still struggle to classify them. In the real world, atomic actions often temporally connect to form more complex composite actions. The challenge lies in recognising composite action of varying durations while other distinct composite or atomic actions occur in the background. Drawing upon the success of relational networks, we propose methods that learn to reason over the semantic concept of objects and actions. We empirically show how ANNs benefit from pretraining, relational inductive biases and unordered set-based latent representations. In this paper we propose deep set conditioned I3D (SCI3D), a two stream relational network that employs latent representation of state and visual representation for reasoning over events and actions. They learn to reason about temporally connected actions in order to identify all of them in the video. The proposed method achieves an improvement of around 1.49% mAP in atomic action recognition and 17.57% mAP in composite action recognition, over a I3D-NL baseline, on the CATER dataset.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionAtomic action recognitionComposite action recognitionTemporal Action Localization

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Atomic action recognition CATER SCI3D Average-mAP 96.77 #1 of 4 Archive leaderboard report
Atomic action recognition CATER R3D-NL Average-mAP 95.28 #2 of 4 Archive leaderboard report
Atomic action recognition CATER Single stream SCI3D Average-mAP 91.82 #3 of 4 Archive leaderboard report
Atomic action recognition CATER FasterRCNN Average-mAP 63.85 #4 of 4 Archive leaderboard report
Composite action recognition CATER Single stream SCI3D Average-mAP 69.76 #1 of 4 Archive leaderboard report
Composite action recognition CATER SCI3D Average-mAP 66.71 #2 of 4 Archive leaderboard report
Composite action recognition CATER R3D-NL Average-mAP 52.19 #3 of 4 Archive leaderboard report
Composite action recognition CATER FasterRCNN Average-mAP 25.45 #4 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections