Papers › VPN: Learning Video-Pose Embedding for Activities of Daily Living

VPN: Learning Video-Pose Embedding for Activities of Daily Living

6 Jul 2020ECCV 2020 8arXiv:2007.03056archive 2025-07-28

Srijan Das, Saurav Sharma, Rui Dai, Francois Bremond, Monique Thonnat

In this paper, we focus on the spatio-temporal aspect of recognizing Activities of Daily Living (ADL). ADL have two specific properties (i) subtle spatio-temporal patterns and (ii) similar visual patterns varying with time. Therefore, ADL may look very similar and often necessitate to look at their fine-grained details to distinguish them. Because the recent spatio-temporal 3D ConvNets are too rigid to capture the subtle visual patterns across an action, we propose a novel Video-Pose Network: VPN. The 2 key components of this VPN are a spatial embedding and an attention network. The spatial embedding projects the 3D poses and RGB cues in a common semantic space. This enables the action recognition framework to learn better spatio-temporal features exploiting both modalities. In order to discriminate similar actions, the attention network provides two functionalities - (i) an end-to-end learnable pose backbone exploiting the topology of human body, and (ii) a coupler to provide joint spatio-temporal attention weights across a video. Experiments show that VPN outperforms the state-of-the-art results for action classification on a large scale human activity dataset: NTU-RGB+D 120, its subset NTU-RGB+D 60, a real-world challenging human activity dataset: Toyota Smarthome and a small scale human-object interaction dataset Northwestern UCLA.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

srijandas07/VPN officialmentioned in papertf report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action ClassificationAction RecognitionHuman-Object Interaction DetectionSkeleton Based Action Recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Classification Toyota Smarthome dataset VPN (RGB + Pose) CS 60.8 #7 of 13 Archive leaderboard report
Action Classification Toyota Smarthome dataset VPN (RGB + Pose) CV1 43.8 #7 of 13 Archive leaderboard report
Action Classification Toyota Smarthome dataset VPN (RGB + Pose) CV2 53.5 #7 of 13 Archive leaderboard report
Action Recognition NTU RGB+D VPN (RGB + Pose) Accuracy (CS) 95.5 #8 of 28 Archive leaderboard report
Action Recognition NTU RGB+D VPN (RGB + Pose) Accuracy (CV) 98.0 #8 of 28 Archive leaderboard report
Action Recognition NTU RGB+D 120 VPN (RGB + Pose) Accuracy (Cross-Setup) 86.3 #16 of 21 Archive leaderboard report
Action Recognition NTU RGB+D 120 VPN (RGB + Pose) Accuracy (Cross-Subject) 87.8 #16 of 21 Archive leaderboard report
Skeleton Based Action Recognition N-UCLA VPN (RGB + Pose) Accuracy 93.5 #19 of 25 Archive leaderboard report
Skeleton Based Action Recognition NTU RGB+D 120 VPN Accuracy (Cross-Setup) 87.8 #40 of 83 Archive leaderboard report
Skeleton Based Action Recognition NTU RGB+D 120 VPN Accuracy (Cross-Subject) 86.3 #40 of 83 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections