{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/boxcars-improving-fine-grained-recognition-of","title":"BoxCars: Improving Fine-Grained Recognition of Vehicles using 3-D Bounding Boxes in Traffic Surveillance","arxiv_id":"1703.00686","date":"2017-03-02","proceeding":null,"authors":["Jakub Sochor","Jakub Špaňhel","Adam Herout"],"abstract":"In this paper, we focus on fine-grained recognition of vehicles mainly in\ntraffic surveillance applications. We propose an approach that is orthogonal to\nrecent advancements in fine-grained recognition (automatic part discovery and\nbilinear pooling). In addition, in contrast to other methods focused on\nfine-grained recognition of vehicles, we do not limit ourselves to a\nfrontal/rear viewpoint, but allow the vehicles to be seen from any viewpoint.\nOur approach is based on 3-D bounding boxes built around the vehicles. The\nbounding box can be automatically constructed from traffic surveillance data.\nFor scenarios where it is not possible to use precise construction, we propose\na method for an estimation of the 3-D bounding box. The 3-D bounding box is\nused to normalize the image viewpoint by \"unpacking\" the image into a plane. We\nalso propose to randomly alter the color of the image and add a rectangle with\nrandom noise to a random position in the image during the training of\nconvolutional neural networks (CNNs). We have collected a large fine-grained\nvehicle data set BoxCars116k, with 116k images of vehicles from various\nviewpoints taken by numerous surveillance cameras. We performed a number of\nexperiments, which show that our proposed method significantly improves CNN\nclassification accuracy (the accuracy is increased by up to 12% points and the\nerror is reduced by up to 50% compared with CNNs without the proposed\nmodifications). We also show that our method outperforms the state-of-the-art\nmethods for fine-grained recognition.","url_abs":"http://arxiv.org/abs/1703.00686v3","url_pdf":"http://arxiv.org/pdf/1703.00686v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"boxcars-improving-fine-grained-recognition-of","repo_url":"https://github.com/JakubSochor/BoxCars","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"3d-object-detection","task_name":"3D Object Detection"},{"task_slug":"vehicle-pose-estimation","task_name":"Vehicle Pose Estimation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1703.00686","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}