Home › Datasets › task › Pose Estimation

Pose Estimation datasets

archive 2025-07-28

124 datasets carry the task tag "Pose Estimation" (the task itself: Pose Estimation), ordered by the archive's paper count. Page 1 of 3: 48 shown of 124. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Pose Estimation datasets 1–48 of 124

The COCO (Common Objects in Context) dataset is a large-scale object detection, segmentation, and captioning dataset.
11,922 papers · 77 benchmarks
KITTI (Karlsruhe Institute of Technology and Toyota Technological Institute) is one of the most popular datasets for use in mobile robotics and autonomous driving.
3,661 papers · 137 benchmarks
MPII (MPII Human Pose)
The MPII Human Pose Dataset for single person pose estimation is composed of about 25K images of which 15K are training samples, 3K are validation samples and 7K are testing samples (which labels are withheld by the authors).
495 papers · 4 benchmarks
The 3D Poses in the Wild dataset is the first dataset in the wild with accurate 3D poses for evaluation.
395 papers · 4 benchmarks
AMASS is a large database of human motion unifying different optical marker-based motion capture datasets by representing them within a common framework and parameterization.
366 papers · 1 benchmark
Description: 10,000 People - Human Pose Recognition Data.
265 papers · 1 benchmark
DensePose (DensePose-COCO)
DensePose-COCO is a large-scale ground-truth dataset with image-to-surface correspondences manually annotated on 50K COCO images and train DensePose-RCNN, to densely regress part-specific UV coordinates within every human region at…
265 papers · 1 benchmark
JHMDB (Joint-annotated Human Motion Data Base)
JHMDB is an action recognition dataset that consists of 960 video sequences belonging to 21 actions.
249 papers · 9 benchmarks
The Pascal3D+ multi-view dataset consists of images in the wild, i.e., images of object categories exhibiting high variability, captured under uncontrolled settings, in cluttered scenes and under many different poses.
237 papers · 1 benchmark
300W (300 Faces-In-The-Wild)
The 300-W is a face dataset that consists of 300 Indoor and 300 Outdoor in-the-wild images.
206 papers · 9 benchmarks
LSP (Leeds Sports Pose)
The Leeds Sports Pose (LSP) dataset is widely used as the benchmark for human pose estimation.
202 papers · 1 benchmark
The YCB-Video dataset is a large-scale video dataset for 6D object pose estimation.
164 papers · 5 benchmarks
MPII Human Pose Dataset is a dataset for human pose estimation.
143 papers · 1 benchmark
The Pix3D dataset is a large-scale benchmark of diverse image-shape pairs with pixel-level 2D-3D alignment.
142 papers · 5 benchmarks
LFPW (Labeled Face Parts in the Wild)
The Labeled Face Parts in-the-Wild (LFPW) consists of 1,432 faces from images downloaded from the web using simple text queries on sites such as google.com, flickr.com, and yahoo.com.
128 papers · 0 benchmarks
SUN3D contains a large-scale RGB-D video database, with 8 annotated sequences.
126 papers · 0 benchmarks
NCLT (North Campus Long-Term Vision and LiDAR)
The NCLT dataset is a large scale, long-term autonomy dataset for robotics research collected on the University of Michigan’s North Campus.
124 papers · 0 benchmarks
The Penn Action Dataset contains 2326 video sequences of 15 different actions and human joint annotations for each sequence.
110 papers · 4 benchmarks
The CrowdPose dataset contains about 20,000 images and a total of 80,000 human poses with 14 labeled keypoints.
99 papers · 2 benchmarks
Aachen Day-Night is a dataset designed for benchmarking 6DOF outdoor visual localization in changing conditions.
93 papers · 1 benchmark
InLoc is a dataset with reference 6DoF poses for large-scale indoor localization.
67 papers · 1 benchmark
This dataset focuses on heavily occluded human with comprehensive annotations including bounding-box, humans pose and instance mask.
66 papers · 6 benchmarks
The TotalCapture dataset consists of 5 subjects performing several activities such as walking, acting, a range of motion sequence (ROM) and freestyle motions, which are recorded using 8 calibrated, static HD RGB cameras and 13 IMUs…
57 papers · 2 benchmarks
The InterHand2.6M dataset is a large-scale real-captured dataset with accurate GT 3D interacting hand poses, used for 3D hand pose estimation The dataset contains 2.6M labeled single and interacting hand frames.
50 papers · 2 benchmarks
UAV-Human is a large dataset for human behavior understanding with UAVs.
47 papers · 5 benchmarks
FLIC (Frames Labelled in Cinema)
The FLIC dataset contains 5003 images from popular Hollywood movies.
46 papers · 2 benchmarks
GazeFollow is a large-scale dataset annotated with the location of where people in images are looking.
38 papers · 1 benchmark
JTA (Joint Track Auto)
JTA is a dataset for people tracking in urban scenarios by exploiting a photorealistic videogame.
36 papers · 1 benchmark
MuCo-3DHP is a large scale training data set showing real images of sophisticated multi-person interactions and occlusions.
35 papers · 0 benchmarks
COCO-WholeBody is an extension of COCO dataset with whole-body annotations.
33 papers · 5 benchmarks
The MannequinChallenge Dataset (MQC) provides in-the-wild videos of people in static poses while a hand-held camera pans around the scene.
28 papers · 0 benchmarks
Animal Kingdom is a large and diverse dataset that provides multiple annotated tasks to enable a more thorough understanding of natural animal behaviors.
26 papers · 2 benchmarks
A three million frame, multi-view, furniture assembly video dataset that includes depth, atomic actions, object segmentation, and human pose.
25 papers · 1 benchmark
The INRIA Person dataset is a dataset of images of persons used for pedestrian detection.
24 papers · 0 benchmarks
KeypointNet is a large-scale and diverse 3D keypoint dataset that contains 83,231 keypoints and 8,329 3D models from 16 object categories, by leveraging numerous human annotations, based on ShapeNet models.
24 papers · 0 benchmarks
ITOP (Invariant-Top View Dataset)
The ITOP dataset consists of 40K training and 10K testing depth images for each of the front-view and top-view tracks.
23 papers · 3 benchmarks
Animal-Pose Dataset is an animal pose dataset to facilitate training and evaluation.
21 papers · 1 benchmark
ModaNet is a street fashion images dataset consisting of annotations related to RGB images.
19 papers · 1 benchmark
A novel dataset facilitating multimodal and Synergetic sociAL Scene Analysis.
18 papers · 1 benchmark
ApolloCar3DT is a dataset that contains 5,277 driving images and over 60K car instances, where each car is fitted with an industry-grade 3D CAD model with absolute model size and semantically labelled keypoints.
17 papers · 14 benchmarks
3D Hand Pose is a multi-view hand pose dataset consisting of color images of hands and different kind of annotations for each: the bounding box and the 2D and 3D location on the joints in the hand.
16 papers · 0 benchmarks
First-Person Hand Action Benchmark is a collection of RGB-D video sequences comprised of more than 100K frames of 45 daily hand action categories, involving 26 different objects in several hand configurations.
15 papers · 2 benchmarks
MSRA Hands is a dataset for hand tracking.
15 papers · 1 benchmark
SVIRO (Synthetic Vehicle Interior Rear Seat Occupancy Dataset)
Contains bounding boxes for object detection, instance segmentation masks, keypoints for pose estimation and depth images for each synthetic scenery as well as images for each individual seat for classification.
15 papers · 0 benchmarks
AIC (AI Challenger)
A large-scale dataset named AIC (AI Challenger) with three sub-datasets, human keypoint detection (HKD), large-scale attribute dataset (LAD) and image Chinese captioning (ICC).
14 papers · 1 benchmark
MVOR (Multi-View Operating Room)
Multi-View Operating Room (MVOR) is a dataset recorded during real clinical interventions.
14 papers · 0 benchmarks
The EgoDexter dataset provides both 2D and 3D pose annotations for 4 testing video sequences with 3190 frames.
13 papers · 0 benchmarks
MoVi (Large Multipurpose Motion and Video Dataset)
Contains 60 female and 30 male actors performing a collection of 20 predefined everyday actions and sports movements, and one self-chosen movement.
12 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.