Datasets › H2O (2 Hands and Objects)

H2O (2 Hands and Objects)

21 Apr 2021 archive 2025-07-28

We present a comprehensive framework for egocentric interaction recognition using markerless 3D annotations of two hands manipulating objects. To this end, we propose a method to create a unified dataset for egocentric 3D interaction recognition. Our method produces annotations of the 3D pose of two hands and the 6D pose of the manipulated objects, along with their interaction labels for each frame. Our dataset, called H2O (2 Hands and Objects), provides synchronized multi-view RGB-D images, interaction labels, object classes, ground-truth 3D poses for left & right hands, 6D object poses, ground-truth camera poses, object meshes and scene point clouds. To the best of our knowledge, this is the first benchmark that enables the study of first-person actions with the use of the pose of both left and right hands manipulating objects and presents an unprecedented level of detail for egocentric 3D interaction recognition. We further propose the method to predict interaction classes by estimating the 3D pose of two hands and the 6D pose of the manipulated objects, jointly from RGB images. Our method models both inter- and intra-dependencies between both hands and objects by learning the topology of a graph convolutional network that predicts interactions. We show that our method facilitated by this dataset establishes a strong baseline for joint hand-object pose estimation and achieves state-of-the-art accuracy for first person interaction recognition.

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Action Recognition H2O (2 Hands and Objects) HandFormer-B/21x8 Actions Top-1 93.39 On the Utility of 3D Hand Poses for Action Recognition s-shamil/HandFormer 11 Compare
Skeleton Based Action Recognition H2O (2 Hands and Objects) CHASE(STSA-Net) Accuracy 94.77 CHASE: Learning Convex Hull Adaptive Shift for... Necolizer/CHASE 4 Compare

Papers archive 2025-07-28

12 shown of 12 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 14. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action Recognition 1 1 9 Oct 2024 ran 2 of 2 samples (0 unverified)
SHARP: Segmentation of Hands and Arms by Range using Pseudo-Depth for Enhanced Egocentric 3D Hand Pose Estimation and Action Recognition 1 2 19 Aug 2024 not harvested
In My Perspective, In My Hands: Accurate Egocentric 2D Hand Pose and Action Recognition 1 2 14 Apr 2024 ran 4 of 5 samples (1 unverified; 5 pointer-only for licence)
On the Utility of 3D Hand Poses for Action Recognition 1 1 14 Mar 2024 ran 8 of 13 samples (5 unverified)
Interactive Spatiotemporal Token Attention Network for Skeleton-based General Interactive Action Recognition 1 2 14 Jul 2023 ran 1 of 1 samples (0 unverified)
Transformer-Based Unified Recognition of Two Hands Manipulating Objects 1 1 1 Jan 2023 not harvested
Hierarchical Temporal Transformer for 3D Hand Pose Estimation and Action Recognition from Egocentric RGB Videos 1 1 20 Sep 2022 ran 1 of 8 samples (7 unverified)
Revisiting Skeleton-based Action Recognition 4 1 28 Apr 2021 not harvested
H2O: Two Hands Manipulating Objects for First Person Interaction Recognition 0 1 22 Apr 2021 not harvested
Disentangling and Unifying Graph Convolutions for Skeleton-Based Action Recognition 3 1 31 Mar 2020 ran 3 of 3 samples (0 unverified)
SlowFast Networks for Video Recognition 15 1 10 Dec 2018 ran 0 of 10 samples (10 unverified)
Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition 24 1 23 Jan 2018 ran 2 of 2 samples (0 unverified)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • H2O (2 Hands and Objects)

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections