Datasets › ITOP
ITOP (Invariant-Top View Dataset)
The ITOP dataset consists of 40K training and 10K testing depth images for each of the front-view and top-view tracks. This dataset contains depth images with 20 actors who perform 15 sequences each and is recorded by two Asus Xtion Pro cameras. The ground-truth of this dataset is the 3D coordinates of 15 body joints.
Source: V2V-PoseNet: Voxel-to-Voxel Prediction Network for Accurate 3D Hand and Human Pose Estimation from a Single Depth Map Image Source: https://www.youtube.com/watch?v=4gPI-GOf9wg
Benchmarks archive 2025-07-28
All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Pose Estimation | ITOP front-view | AdaPose Mean mAP 93.38 | Sequential 3D Human Pose Estimation Using Adaptive Point... | Hmslab/Adapose | 7 | Compare |
| Pose Estimation | ITOP top-view | DECA-D3 Mean mAP 86.92 | DECA: Deep viewpoint-Equivariant human pose estimation... | mmlab-cv/deca | 5 | Compare |
| 3D Human Pose Estimation | ITOP front-view | SPiKE Mean mAP 89.19 | SPiKE: 3D Human Pose from Point Cloud Sequences | iballester/SPiKE | 1 | Compare |
Papers archive 2025-07-28
7 shown of 7 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 23. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| SPiKE: 3D Human Pose from Point Cloud Sequences | 1 | 2 | 3 Sep 2024 | not harvested |
| Sequential 3D Human Pose Estimation Using Adaptive Point Cloud Sampling Strategy | 1 | 1 | 19 Aug 2021 | not harvested |
| DECA: Deep viewpoint-Equivariant human pose estimation using Capsule Autoencoders | 1 | 3 | 19 Aug 2021 | not harvested |
| A2J: Anchor-to-Joint Regression Network for 3D Articulated Pose Estimation from a Single Depth Image | 2 | 1 | 27 Aug 2019 | ran 6 of 9 samples (3 unverified) |
| V2V-PoseNet: Voxel-to-Voxel Prediction Network for Accurate 3D Hand and Human Pose Estimation from a Single Depth Map | 5 | 2 | 20 Nov 2017 | not harvested |
| Towards Good Practices for Deep 3D Hand Pose Estimation | 0 | 2 | 23 Jul 2017 | not harvested |
| Towards Viewpoint Invariant 3D Human Pose Estimation | 2 | 2 | 23 Mar 2016 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- ITOP front-view
- ITOP top-view
- ITOP front-view
- ITOP
4 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections