Papers › Transformer-Based Unified Recognition of Two Hands Manipulating Objects
Transformer-Based Unified Recognition of Two Hands Manipulating Objects
Hoseong Cho, Chanwoo Kim, Jihyeon Kim, Seongyeong Lee, Elkhan Ismayilzada, Seungryul Baek
Understanding the hand-object interactions from an egocentric video has received a great attention recently. So far, most approaches are based on the convolutional neural network (CNN) features combined with the temporal encoding via the long short-term memory (LSTM) or graph convolution network (GCN) to provide the unified understanding of two hands, an object and their interactions. In this paper, we propose the Transformer-based unified framework that provides better understanding of two hands manipulating objects. In our framework, we insert the whole image depicting two hands, an object and their interactions as input and jointly estimate 3 information from each frame: poses of two hands, pose of an object and object types. Afterwards, the action class defined by the hand-object interactions is predicted from the entire video based on the estimated information combined with the contact map that encodes the interaction between two hands and an object. Experiments are conducted on H2O and FPHA benchmark datasets and we demonstrated the superiority of our method achieving the state-of-the-art accuracy. Ablative studies further demonstrate the effectiveness of each proposed module.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Action Recognition | H2O (2 Hands and Objects) | H2OTR | Actions Top-1 | 90.90 | #4 of 11 | Archive leaderboard | report |
| Action Recognition | H2O (2 Hands and Objects) | H2OTR | Hand Pose | 3D (est.) | #4 of 11 | Archive leaderboard | report |
| Action Recognition | H2O (2 Hands and Objects) | H2OTR | Object Label | Yes (est.) | #4 of 11 | Archive leaderboard | report |
| Action Recognition | H2O (2 Hands and Objects) | H2OTR | Object Pose | Yes (est.) | #4 of 11 | Archive leaderboard | report |
| Action Recognition | H2O (2 Hands and Objects) | H2OTR | RGB | Yes | #4 of 11 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections