{"url":"/dataset/cityscapes-3d","name":"Cityscapes 3D","full_name":null,"description_markdown":"Detecting vehicles and representing their position and orientation in the three dimensional space is a key technology for autonomous driving. Recently, methods for 3D vehicle detection solely based on monocular RGB images gained popularity. In order to facilitate this task as well as to compare and drive state-of-the-art methods, several new datasets and benchmarks have been published. Ground truth annotations of vehicles are usually obtained using lidar point clouds, which often induces errors due to imperfect calibration or synchronization between both sensors. To this end, we propose Cityscapes 3D, extending the original Cityscapes dataset with 3D bounding box annotations for all types of vehicles. In contrast to existing datasets, our 3D annotations were labeled using stereo RGB images only and capture all nine degrees of freedom. This leads to a pixel-accurate reprojection in the RGB image and a higher range of annotations compared to lidar-based approaches. In order to ease multitask learning, we provide a pairing of 2D instance segments with 3D bounding boxes. In addition, we complement the Cityscapes benchmark suite with 3D vehicle detection based on the new annotations as well as metrics presented in this work. Dataset and benchmark are available online.","description_withheld":null,"homepage":"https://github.com/mcordts/cityscapesScripts","introduced_date":"2020-06-14","introduced_date_note":null,"introduced_by":{"paper":"/paper/cityscapes-3d-dataset-and-benchmark-for-9-dof","title":"Cityscapes 3D: Dataset and Benchmark for 9 DoF Vehicle Detection","first_author":"Nils Gählert","url":null},"license":{"name":"MIT license","url":"https://github.com/mcordts/cityscapesScripts/blob/master/LICENSE"},"modalities":[{"name":"Images","url":"/datasets/modality/images"}],"tasks":[{"name":"Semantic Segmentation","url":"/task/semantic-segmentation","datasets_with_task":"/datasets/task/semantic-segmentation"},{"name":"3D Object Detection","url":"/task/3d-object-detection","datasets_with_task":"/datasets/task/3d-object-detection"},{"name":"Monocular Depth Estimation","url":"/task/monocular-depth-estimation","datasets_with_task":"/datasets/task/monocular-depth-estimation"}],"languages":[],"variants":["Cityscapes 3D"],"data_loaders":[],"num_papers_in_archive":13,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/3d-object-detection-on-cityscapes-3d","task":"3D Object Detection","dataset_variant":"Cityscapes 3D","rows":1,"metrics":["mDS"],"first_row_in_archive_order":{"model":"TaskPrompter","paper":"/paper/joint-2d-3d-multi-task-learning-on-cityscapes","metrics":{"mDS":"32.94"},"code_links":[{"title":"prismformore/multi-task-transformer","url":"https://github.com/prismformore/multi-task-transformer"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/monocular-depth-estimation-on-cityscapes-3d","task":"Monocular Depth Estimation","dataset_variant":"Cityscapes 3D","rows":1,"metrics":["RMSE"],"first_row_in_archive_order":{"model":"TaskPrompter","paper":"/paper/joint-2d-3d-multi-task-learning-on-cityscapes","metrics":{"RMSE":"6.78"},"code_links":[{"title":"prismformore/multi-task-transformer","url":"https://github.com/prismformore/multi-task-transformer"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/semantic-segmentation-on-cityscapes-3d","task":"Semantic Segmentation","dataset_variant":"Cityscapes 3D","rows":1,"metrics":["mIoU"],"first_row_in_archive_order":{"model":"TaskPrompter","paper":"/paper/joint-2d-3d-multi-task-learning-on-cityscapes","metrics":{"mIoU":"77.72"},"code_links":[{"title":"prismformore/multi-task-transformer","url":"https://github.com/prismformore/multi-task-transformer"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/joint-2d-3d-multi-task-learning-on-cityscapes","title":"Joint 2D-3D Multi-Task Learning on Cityscapes-3D: 3D Detection, Segmentation, and Depth Estimation","date":"2023-04-03","rows_on_this_dataset":3,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}