Home › Datasets › task › Scene Understanding
Scene Understanding datasets
archive 2025-07-28
43 datasets carry the task tag "Scene Understanding" (the task itself: Scene Understanding), ordered by the archive's paper count. Page 1 of 1: 43 shown of 43. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Scene Understanding datasets 1–43 of 43
The ADE20K semantic segmentation dataset contains more than 20K scene-centric images exhaustively annotated with pixel-level objects and object parts labels.
1,213 papers · 32 benchmarks
RELLIS-3D is a multi-modal dataset for off-road robotics.
57 papers · 3 benchmarks
UAVid is a high-resolution UAV semantic segmentation dataset as a complement, which brings new challenges, including large scale variation, moving object recognition and temporal consistency preservation.
54 papers · 2 benchmarks
The large-scale MUSIC-AVQA dataset of musical performance contains 45,867 question-answer pairs, distributed in 9,288 videos for over 150 hours.
51 papers · 1 benchmark
A novel dataset and benchmark, which features 1482 RGB-D scans of 478 environments across multiple time steps.
44 papers · 4 benchmarks
KITTI Road is road and lane estimation benchmark that consists of 289 training and 290 test images.
41 papers · 0 benchmarks
SUIM (Segmentation of Underwater IMagery)
The Segmentation of Underwater IMagery (SUIM) dataset contains over 1500 images with pixel annotations for eight object categories: fish (vertebrates), reefs (invertebrates), aquatic plants, wrecks/ruins, human divers, robots, and…
34 papers · 2 benchmarks
The SensatUrbat dataset is an urban-scale photogrammetric point cloud dataset with nearly three billion richly annotated points, which is five times the number of labeled points than the existing largest point cloud dataset.
28 papers · 1 benchmark
RADIATE (RAdar Dataset In Adverse weaThEr)
RADIATE (RAdar Dataset In Adverse weaThEr) is new automotive dataset created by Heriot-Watt University which includes Radar, Lidar, Stereo Camera and GPS/IMU.
24 papers · 2 benchmarks
Toronto-3D is a large-scale urban outdoor point cloud dataset acquired by an MLS system in Toronto, Canada for semantic segmentation.
24 papers · 2 benchmarks
A simulation-based dataset featuring 20,000 stack configurations composed of a variety of elementary geometric primitives richly annotated regarding semantics and structural stability.
22 papers · 2 benchmarks
MLRSNet is a a multi-label high spatial resolution remote sensing dataset for semantic scene understanding.
17 papers · 2 benchmarks
SceneNet is a dataset of labelled synthetic indoor scenes.
17 papers · 0 benchmarks
DADA-2000 is a large-scale benchmark with 2000 video sequences (named as DADA-2000) is contributed with laborious annotation for driver attention (fixation, saccade, focusing time), accident objects/intervals, as well as the accident…
16 papers · 0 benchmarks
A novel benchmark dataset that includes a manually annotated point cloud for over 260 million laser scanning points into 100'000 (approx.) assets from Dublin LiDAR point cloud [12] in 2015.
15 papers · 0 benchmarks
IRS (Indoor Robotics Stereo)
IRS is an open dataset for indoor robotics vision tasks, especially disparity and surface normal estimation.
12 papers · 0 benchmarks
DAWN emphasizes a diverse traffic environment (urban, highway and freeway) as well as a rich variety of traffic flow.
10 papers · 0 benchmarks
DeepScores contains high quality images of musical scores, partitioned into 300,000 sheets of written music that contain symbols of different shapes and sizes.
10 papers · 0 benchmarks
The AIRS (Aerial Imagery for Roof Segmentation) dataset provides a wide coverage of aerial imagery with 7.5 cm resolution and contains over 220,000 buildings.
9 papers · 1 benchmark
PSI-AVA is a dataset designed for holistic surgical scene understanding.
9 papers · 0 benchmarks
The Pascal Panoptic Parts dataset consists of annotations for the part-aware panoptic segmentation task on the PASCAL VOC 2010 dataset.
9 papers · 2 benchmarks
The Cityscapes Panoptic Parts dataset introduces part-aware panoptic segmentation annotations for the Cityscapes dataset.
8 papers · 1 benchmark
DeepLoc is a large-scale urban outdoor localization dataset.
8 papers · 0 benchmarks
SynPick is a synthetic dataset for dynamic scene understanding in bin-picking scenarios.
7 papers · 1 benchmark
A Large Dataset of Object Scans is a dataset of more than ten thousand 3D scans of real objects.
5 papers · 0 benchmarks
A new RGB-D video dataset, i.e., UCLA Human-Human-Object Interaction (HHOI) dataset, which includes 3 types of human-human interactions, i.e., shake hands, high-five, pull up, and 2 types of human-object-human interactions, i.e., throw and…
5 papers · 0 benchmarks
The semantic line (SEL) dataset contains 1,750 outdoor images in total, which are split into 1,575 training and 175 testing images.
5 papers · 1 benchmark
WWW Crowd provides 10,000 videos with over 8 million frames from 8,257 diverse scenes, therefore offering a comprehensive dataset for the area of crowd understanding.
5 papers · 0 benchmarks
AeroRIT is a hyperspectral dataset to facilitate aerial hyperspectral scene understanding.
4 papers · 0 benchmarks
CDS2K is a benchmark for Concealed scene understanding (CSU), which is a hot computer vision topic aiming to perceive objects with camouflaged properties.
4 papers · 0 benchmarks
The VideoNavQA dataset contains pairs of questions and videos generated in the House3D environment.
4 papers · 0 benchmarks
A new resource to train and evaluate multitask systems on samples in multiple modalities and three languages.
3 papers · 0 benchmarks
Includes challenging sequences and extensive data stratification in-terms of camera and object motion, velocity magnitudes, direction, and rotational speeds.
3 papers · 0 benchmarks
The RailEye3D dataset, a collection of train-platform scenarios for applications targeting passenger safety and automation of train dispatching, consists of 10 image sequences captured at 6 railway stations in Austria.
3 papers · 0 benchmarks
InstaOrder can be used to understand the geometrical relationships of instances in an image.
2 papers · 0 benchmarks
Stanford-ECM is an egocentric multimodal dataset which comprises about 27 hours of egocentric video augmented with heart rate and acceleration data.
2 papers · 0 benchmarks
TDW is a 3D virtual world simulation platform, utilizing state-of-the-art video game engine technology.
1 paper · 0 benchmarks
The Apron Dataset focuses on training and evaluating classification and detection models for airport-apron logistics.
1 paper · 0 benchmarks
COQE (Containers Of liQuid contEnt)
Contains more than 5,000 images of 10,000 liquid containers in context labelled with volume, amount of content, bounding box annotation, and corresponding similar 3D CAD models.
1 paper · 0 benchmarks
DeepLocCross is a localization dataset that contains RGB-D stereo images captured at 1280 x 720 pixels at a rate of 20 Hz.
1 paper · 0 benchmarks
The dfdindoor dataset contains 110 images for training and 29 images for testing.
1 paper · 0 benchmarks
NavigationNet is a computer vision dataset and benchmark to allow the utilization of deep reinforcement learning on scene-understanding-based indoor navigation.
1 paper · 0 benchmarks
Panoramic Video Panoptic Segmentation Dataset is a large-scale dataset that offers high-quality panoptic segmentation labels for autonomous driving.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.