{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/apollocar3d-a-large-3d-car-instance","title":"ApolloCar3D: A Large 3D Car Instance Understanding Benchmark for Autonomous Driving","arxiv_id":"1811.12222","date":"2018-11-29","proceeding":"CVPR 2019 6","authors":["Xibin Song","Peng Wang","Dingfu Zhou","Rui Zhu","Chenye Guan","Yuchao Dai","Hao Su","Hongdong Li","Ruigang Yang"],"abstract":"Autonomous driving has attracted remarkable attention from both industry and\nacademia. An important task is to estimate 3D properties(e.g.translation,\nrotation and shape) of a moving or parked vehicle on the road. This task, while\ncritical, is still under-researched in the computer vision community -\npartially owing to the lack of large scale and fully-annotated 3D car database\nsuitable for autonomous driving research. In this paper, we contribute the\nfirst large-scale database suitable for 3D car instance understanding -\nApolloCar3D. The dataset contains 5,277 driving images and over 60K car\ninstances, where each car is fitted with an industry-grade 3D CAD model with\nabsolute model size and semantically labelled keypoints. This dataset is above\n20 times larger than PASCAL3D+ and KITTI, the current state-of-the-art. To\nenable efficient labelling in 3D, we build a pipeline by considering 2D-3D\nkeypoint correspondences for a single instance and 3D relationship among\nmultiple instances. Equipped with such dataset, we build various baseline\nalgorithms with the state-of-the-art deep convolutional neural networks.\nSpecifically, we first segment each car with a pre-trained Mask R-CNN, and then\nregress towards its 3D pose and shape based on a deformable 3D car model with\nor without using semantic keypoints. We show that using keypoints significantly\nimproves fitting performance. Finally, we develop a new 3D metric jointly\nconsidering 3D pose and 3D shape, allowing for comprehensive evaluation and\nablation study. By comparing with human performance we suggest several future\ndirections for further improvements.","url_abs":"http://arxiv.org/abs/1811.12222v2","url_pdf":"http://arxiv.org/pdf/1811.12222v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"3d-car-instance-understanding","task_name":"3D Car Instance Understanding"},{"task_slug":"autonomous-driving","task_name":"Autonomous Driving"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"mask-r-cnn","method_name":"Mask R-CNN"},{"method_slug":"rpn","method_name":"RPN"},{"method_slug":"roi-align","method_name":"RoIAlign"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[{"slug":"apollocar3d","name":"ApolloCar3D","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1811.12222","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}