Papers › Transformer-Based Unified Recognition of Two Hands Manipulating Objects

Transformer-Based Unified Recognition of Two Hands Manipulating Objects

1 Jan 2023CVPR 2023 1archive 2025-07-28

Hoseong Cho, Chanwoo Kim, Jihyeon Kim, Seongyeong Lee, Elkhan Ismayilzada, Seungryul Baek

Understanding the hand-object interactions from an egocentric video has received a great attention recently. So far, most approaches are based on the convolutional neural network (CNN) features combined with the temporal encoding via the long short-term memory (LSTM) or graph convolution network (GCN) to provide the unified understanding of two hands, an object and their interactions. In this paper, we propose the Transformer-based unified framework that provides better understanding of two hands manipulating objects. In our framework, we insert the whole image depicting two hands, an object and their interactions as input and jointly estimate 3 information from each frame: poses of two hands, pose of an object and object types. Afterwards, the action class defined by the hand-object interactions is predicted from the entire video based on the estimated information combined with the contact map that encodes the interaction between two hands and an object. Experiments are conducted on H2O and FPHA benchmark datasets and we demonstrated the superiority of our method achieving the state-of-the-art accuracy. Ablative studies further demonstrate the effectiveness of each proposed module.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action RecognitionObject

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Action Recognition H2O (2 Hands and Objects) H2OTR Actions Top-1 90.90 #4 of 11 Archive leaderboard report
Action Recognition H2O (2 Hands and Objects) H2OTR Hand Pose 3D (est.) #4 of 11 Archive leaderboard report
Action Recognition H2O (2 Hands and Objects) H2OTR Object Label Yes (est.) #4 of 11 Archive leaderboard report
Action Recognition H2O (2 Hands and Objects) H2OTR Object Pose Yes (est.) #4 of 11 Archive leaderboard report
Action Recognition H2O (2 Hands and Objects) H2OTR RGB Yes #4 of 11 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Convolution

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections