Home › Datasets › modality › 3D

3D datasets

archive 2025-07-28

380 datasets carry the modality tag "3D", ordered by the archive's paper count. Page 1 of 8: 48 shown of 380. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

3D datasets 1–48 of 380

ShapeNet is a large scale repository for 3D CAD models developed by researchers from Stanford University, Princeton University and the Toyota Technological Institute at Chicago, USA.
1,947 papers · 13 benchmarks
S3DIS (Stanford 3D Indoor Scene Dataset (S3DIS))
The Stanford 3D Indoor Scene Dataset (S3DIS) dataset contains 6 large-scale indoor areas with 271 rooms.
488 papers · 9 benchmarks
The Waymo Open Dataset is comprised of high resolution sensor data collected by autonomous vehicles operated by the Waymo Driver in a wide variety of conditions.
481 papers · 16 benchmarks
The Matterport3D dataset is a large RGB-D dataset for scene understanding in indoor environments.
461 papers · 4 benchmarks
The Replica Dataset is a dataset of high quality reconstructions of a variety of indoor spaces.
414 papers · 4 benchmarks
The 3D Poses in the Wild dataset is the first dataset in the wild with accurate 3D poses for evaluation.
395 papers · 4 benchmarks
Objaverse is a large dataset of objects with 800K+ (and growing) 3D models with descriptive captions, tags, and animations.
393 papers · 2 benchmarks
AMASS is a large database of human motion unifying different optical marker-based motion capture datasets by representing them within a common framework and parameterization.
366 papers · 1 benchmark
MPI-INF-3DHP is a 3D human body pose estimation dataset consisting of both constrained indoor and complex outdoor scenes.
289 papers · 5 benchmarks
The Pascal3D+ multi-view dataset consists of images in the wild, i.e., images of object categories exhibiting high variability, captured under uncontrolled settings, in cluttered scenes and under many different poses.
237 papers · 1 benchmark
HumanML3D is a 3D human motion-language dataset that originates from a combination of HumanAct12 and Amass dataset.
201 papers · 2 benchmarks
SUNCG is a large-scale dataset of synthetic 3D scenes with dense volumetric annotations.
186 papers · 0 benchmarks
ShapeNetCore is a subset of the full ShapeNet dataset with single clean 3D models and manually verified category and alignment annotations.
180 papers · 1 benchmark
PartNet is a consistent, large-scale dataset of 3D objects annotated with fine-grained, instance-level, and hierarchical 3D part information.
156 papers · 3 benchmarks
The Pix3D dataset is a large-scale benchmark of diverse image-shape pairs with pixel-level 2D-3D alignment.
142 papers · 5 benchmarks
SUN3D contains a large-scale RGB-D video database, with 8 annotated sequences.
126 papers · 0 benchmarks
FreiHAND is a 3D hand pose dataset which records different hand actions performed by 32 people.
125 papers · 1 benchmark
AFLW2000-3D is a dataset of 2000 images that have been annotated with image-level 68-point 3D facial landmarks.
117 papers · 8 benchmarks
The smallNORB dataset is a datset for 3D object recognition from shape.
112 papers · 1 benchmark
The Chairs dataset contains rendered images of around 1000 different three-dimensional chair models.
109 papers · 1 benchmark
Despite the considerable progress in automatic abdominal multi-organ segmentation from CT/MRI scans in recent years, a comprehensive evaluation of the models' capabilities is hampered by the lack of a large-scale benchmark from diverse…
105 papers · 1 benchmark
The ABC Dataset is a collection of one million Computer-Aided Design (CAD) models for research of geometric deep learning methods and applications.
104 papers · 0 benchmarks
The BP4D-Spontaneous dataset is a 3D video database of spontaneous facial expressions in a diverse group of young adults.
104 papers · 3 benchmarks
ReferIt3D provides two large-scale and complementary visio-linguistic datasets: i) Sr3D, which contains 83.5K template-based utterances leveraging spatial relations among fine-grained object classes to localize a referred object in a…
104 papers · 1 benchmark
HM3D (Habitat-Matterport 3D)
Habitat-Matterport 3D (HM3D) is a large-scale dataset of 1,000 building-scale 3D reconstructions from a diverse set of real-world locations.
3D
96 papers · 0 benchmarks
T-LESS is a dataset for estimating the 6D pose, i.e.
94 papers · 2 benchmarks
Aachen Day-Night is a dataset designed for benchmarking 6DOF outdoor visual localization in changing conditions.
93 papers · 1 benchmark
FaceWarehouse is a 3D facial expression database that provides the facial geometry of 150 subjects, covering a wide range of ages and ethnic backgrounds.
91 papers · 0 benchmarks
Dataset produced for the SAPIEN simulation environment.
88 papers · 0 benchmarks
Structured3D is a large-scale photo-realistic dataset containing 3.5K house designs (a) created by professional designers with a variety of ground truth 3D structure annotations (b) and generate photo-realistic 2D images (c).
85 papers · 7 benchmarks
ABO (Amazon Berkeley Objects)
ABO is a large-scale dataset designed for material prediction and multi-view retrieval experiments.
82 papers · 0 benchmarks
CoMA contains 17,794 meshes of the human face in various expressions Source: DEMEA: Deep Mesh Autoencoders for Non-Rigidly Deforming Objects Image Source: https://coma.is.tue.mpg.de/
81 papers · 1 benchmark
BABEL is a large dataset with language labels describing the actions being performed in mocap sequences.
72 papers · 1 benchmark
Scan2CAD is an alignment dataset based on 1506 ScanNet scans with 97607 annotated keypoints pairs between 14225 (3049 unique) CAD models from ShapeNet and their counterpart objects in the scans.
72 papers · 1 benchmark
The SemanticPOSS dataset for 3D semantic segmentation contains 2988 various and complicated LiDAR scans with large quantity of dynamic instances.
71 papers · 1 benchmark
MuPoTS-3D (Multiperson Pose Test Set in 3DMulti-person Pose estimation Test Set in 3D)
MuPoTs-3D (Multi-person Pose estimation Test Set in 3D) is a dataset for pose estimation composed of more than 8,000 frames from 20 real-world scenes with up to three subjects.
70 papers · 3 benchmarks
OmniObject3D is a large vocabulary 3D object dataset with massive high-quality real-scanned 3D objects.
70 papers · 0 benchmarks
AGORA is a synthetic human dataset with high realism and accurate ground truth.
68 papers · 4 benchmarks
Semantic3D is a point cloud dataset of scanned outdoor scenes with over 3 billion points.
66 papers · 1 benchmark
SceneNN is an RGB-D scene dataset consisting of more than 100 indoor scenes.
63 papers · 1 benchmark
VOCASET is a 4D face dataset with about 29 minutes of 4D scans captured at 60 fps and synchronized audio.
60 papers · 1 benchmark
SQA3D (Situated Question Answering in 3D Scenes)
SQA3D is a dataset for embodied scene understanding, where an agent needs to comprehend the scene it situates from an first person's perspective and answer questions.
58 papers · 3 benchmarks
The TotalCapture dataset consists of 5 subjects performing several activities such as walking, acting, a range of motion sequence (ROM) and freestyle motions, which are recorded using 8 calibrated, static HD RGB cameras and 13 IMUs…
57 papers · 2 benchmarks
BEAT (Body-Expression-Audio-Text)
BEAT has i) 76 hours, high-quality, multi-modal data captured from 30 speakers talking with eight different emotions and in four different languages, ii) 32 millions frame-level emotion and semantic relevance annotations.
56 papers · 1 benchmark
The SEMAINE videos dataset contains spontaneous data capturing the audiovisual interaction between a human and an operator undertaking the role of an avatar with four personalities: Poppy (happy), Obadiah (gloomy), Spike (angry) and…
54 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.