Home › Datasets › task › Pose Estimation
Pose Estimation datasets
archive 2025-07-28
124 datasets carry the task tag "Pose Estimation" (the task itself: Pose Estimation), ordered by the archive's paper count. Page 2 of 3: 48 shown of 124. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Pose Estimation datasets 49–96 of 124
SLOPER4D is a novel scene-aware dataset collected in large urban environments to facilitate the research of global human pose estimation (GHPE) with human-scene interaction in the wild.
11 papers · 1 benchmark
UnrealEgo is a dataset that provides in-the-wild stereo images with a large variety of motions for 3D human pose estimation.
11 papers · 1 benchmark
Unite The People is a dataset for 3D body estimation.
10 papers · 0 benchmarks
xR-EgoPose is an egocentric synthetic dataset for egocentric 3D human pose estimation.
10 papers · 0 benchmarks
We provide manual annotations of 14 semantic keypoints for 100,000 car instances (sedan, suv, bus, and truck) from 53,000 images captured from 18 moving cameras at Multiple intersections in Pittsburgh, PA.
9 papers · 2 benchmarks
EgoCap is a dataest of 100,000 egocentric images of eight people in different clothing, with 75,000 images from six people used for training.
9 papers · 0 benchmarks
HUMAN4D is a large and multimodal 4D dataset that contains a variety of human activities simultaneously captured by a professional marker-based MoCap, a volumetric capture and an audio recording system.
9 papers · 0 benchmarks
ATRW (Amur Tiger Re-identification in the Wild)
The ATRW Dataset contains over 8,000 video clips from 92 Amur tigers, with bounding box, pose keypoint, and tiger identity annotations.
8 papers · 0 benchmarks
GPA (Geometric Pose Affordance)
multi-view imagery of people interacting with a variety of rich 3D environments
8 papers · 2 benchmarks
H3WB (Human 3.6M 3D WholeBody)
Human3.6M 3D WholeBody (H3WB) is a large scale dataset with 133 whole-body keypoint annotations on 100K images, made possible by a new multi-view pipeline.
8 papers · 3 benchmarks
The IMUPoser Dataset is a dataset for estimating body pose using IMUs already in devices that many users own -- namely smartphones, smartwatches, and earbuds.
8 papers · 0 benchmarks
Biwi Kinect Head Pose is a challenging dataset mainly inspired by the automotive setup.
7 papers · 0 benchmarks
The HandNet dataset contains depth images of 10 participants' hands non-rigidly deforming in front of a RealSense RGB-D camera.
7 papers · 0 benchmarks
The odometry benchmark consists of 22 stereo sequences, saved in loss less png format: We provide 11 sequences (00-10) with ground truth trajectories for training and 11 sequences (11-21) without ground truth for evaluation.
7 papers · 1 benchmark
MonoPerfCap is a benchmark dataset for human 3D performance capture from monocular video input consisting of around 40k frames, which covers a variety of different scenarios.
7 papers · 0 benchmarks
The SynthHands dataset is a dataset for hand pose estimation which consists of real captured hand motion retargeted to a virtual hand with natural backgrounds and interactions with different objects.
7 papers · 0 benchmarks
Falling Things (FAT) is a dataset for advancing the state-of-the-art in object detection and 3D pose estimation in the context of robotics.
6 papers · 0 benchmarks
A new dataset with significant occlusions related to object manipulation.
6 papers · 0 benchmarks
Dataset for large-scale yoga pose recognition with 82 classes.
6 papers · 0 benchmarks
BRACE (The Breakdancing Competition Dataset for Dance Motion Synthesis)
BRACE is a dataset for audio-conditioned dance motion synthesis challenging common assumptions for this task: - strong music-dance correlation - controlled motion data - simple poses and movements To address these issues: - We focus on…
5 papers · 2 benchmarks
DREAM-dataset (Deep Robot-to-camera Extrinsics for Articulated Manipulators)
The DREAM dataset is introduce by the paper "Camera-to-Robot Pose Estimation from a Single Image" (ICRA 2020).
5 papers · 1 benchmark
PedX is a large-scale multi-modal collection of pedestrians at complex urban intersections.
5 papers · 0 benchmarks
FewSOL (A Dataset for Few-Shot Object Learning in Robotic Environments)
The Few-Shot Object Learning (FewSOL) dataset can be used for object recognition with a few images per object.
4 papers · 0 benchmarks
The Fraunhofer IPA Bin-Picking dataset is a large-scale dataset comprising both simulated and real-world scenes for various objects (potentially having symmetries) and is fully annotated with 6D poses.
4 papers · 0 benchmarks
HuPR (Human Pose with Millimeter Wave Radar)
HuPR is a human pose estimation benchmark is created using cross-calibrated mmWave radar sensors and a monocular RGB camera for cross-modality training of radar-based human pose estimation.
4 papers · 0 benchmarks
MERL-RAV (MERL Reannotation of AFLW with Visibility)
The MERL-RAV (MERL Reannotation of AFLW with Visibility) Dataset contains over 19,000 face images in a full range of head poses.
4 papers · 2 benchmarks
SportsPose (SportsPose - A Dynamic 3D sports pose dataset)
Accurate 3D human pose estimation is essential for sports analytics, coaching, and injury prevention.
4 papers · 0 benchmarks
Amateur Drawings is a dataset collected via the public demo of Animated Drawings, containing over 178,000 amateur drawings and corresponding user-accepted character bounding boxes, segmentation masks, and joint location annotations.
3 papers · 0 benchmarks
CHAIRS is a large-scale motion-captured f-AHOI dataset, consisting of 17.3 hours of versatile interactions between 46 participants and 81 articulated and rigid sittable objects.
3 papers · 0 benchmarks
CORSMAL is a dataset for estimating the position and orientation in 3D (or 6D pose) of an object from a single view.
3 papers · 0 benchmarks
The Composable activities dataset consists of 693 videos that contain activities in 16 classes performed by 14 actors.
3 papers · 0 benchmarks
Fitness-AQA (Fitness Action Quality Assessment [ECCV 2022])
Largest, first-of-its-kind, in-the-wild, fine-grained workout/exercise posture analysis dataset, covering three different exercises: BackSquat, Barbell Row, and Overhead Press.
3 papers · 0 benchmarks
HOPE-Image (Household Objects for Pose Estimation)
The NVIDIA HOPE datasets consist of RGBD images and video sequences with labeled 6-DoF poses for 28 toy grocery objects.
3 papers · 0 benchmarks
HOPE-Video (Household Objects for Pose Estimation)
The HOPE-Video dataset contains 10 video sequences (2038 frames) with 5-20 objects on a tabletop scene captured by a robot arm-mounted RealSense D415 RGBD camera.
3 papers · 0 benchmarks
The ICVL dataset is a hand pose estimation dataset that consists of 330K training frames and 2 testing sequences with each 800 frames.
3 papers · 0 benchmarks
PoPArt (Poses of People in Art: A Data Set for Human Pose Estimation in Digital Art History)
Throughout the history of art, the pose—as the holistic abstraction of the human body's expression—has proven to be a constant in numerous studies.
3 papers · 1 benchmark
VBR (VBR: A Vision Benchmark in Rome)
This dataset presents a vision and perception research dataset collected in Rome, featuring RGB data, 3D point clouds, IMU, and GPS data.
3 papers · 0 benchmarks
BASEPROD (The Bardenas Semi-Desert Planetary Rover Dataset)
BASEPROD provides comprehensive rover sensor data collected over a 1.7 km traverse, accompanied by high-resolution 2D and 3D drone maps of the terrain.
2 papers · 0 benchmarks
Estimating camera motion in deformable scenes poses a complex and open research challenge.
2 papers · 1 benchmark
Dataset page: https://github.com/mosamdabhi/MBW-Data MBW - Zoo is a challenging dataset consisting image frames of tail-end distribution categories (such as Fish, Colobus Monkeys, Chimpanzees, etc.) with their corresponding 2D, 3D, and…
2 papers · 0 benchmarks
MOTFront provides photo-realistic RGB-D images with their corresponding instance segmentation masks, class labels, 2D & 3D bounding boxes, 3D geometry, 3D poses and camera parameters.
2 papers · 0 benchmarks
MPHOI-72 (Multi-person Human-object Interaction Dataset 72)
MPHOI-72 is a multi-person human-object interaction dataset that can be used for a wide variety of HOI/activity recognition and pose estimation/object tracking tasks.
2 papers · 0 benchmarks
The data includes all movement trajectories extracted from the videos of Parkinson's assessments using Convolutional Pose Machines (CPM) as well as the confidence values from CPM.
2 papers · 0 benchmarks
The Poser dataset is a dataset for pose estimation which consists of 1927 training and 418 test images.
2 papers · 0 benchmarks
Rendered Handpose Dataset contains 41258 training and 2728 testing samples.
2 papers · 0 benchmarks
The Retinal Microsurgery dataset is a dataset for surgical instrument tracking.
2 papers · 0 benchmarks
This dataset comprehends the 3D building information model (in IFC and Revit formats), manually elaborated based on the terrestrial laser scanner of the sequence 2 of ConSLAM, and the refined ground truth (GT) poses (in TUM format) of…
2 papers · 0 benchmarks
~6 million synthetic depth frames for pose estimation from multiple cameras.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.