{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-simple-baseline-for-bev-perception-without","title":"Simple-BEV: What Really Matters for Multi-Sensor BEV Perception?","arxiv_id":"2206.07959","date":"2022-06-16","proceeding":null,"authors":["Adam W. Harley","Zhaoyuan Fang","Jie Li","Rares Ambrus","Katerina Fragkiadaki"],"abstract":"Building 3D perception systems for autonomous vehicles that do not rely on high-density LiDAR is a critical research problem because of the expense of LiDAR systems compared to cameras and other sensors. Recent research has developed a variety of camera-only methods, where features are differentiably \"lifted\" from the multi-camera images onto the 2D ground plane, yielding a \"bird's eye view\" (BEV) feature representation of the 3D space around the vehicle. This line of work has produced a variety of novel \"lifting\" methods, but we observe that other details in the training setups have shifted at the same time, making it unclear what really matters in top-performing methods. We also observe that using cameras alone is not a real-world constraint, considering that additional sensors like radar have been integrated into real vehicles for years already. In this paper, we first of all attempt to elucidate the high-impact factors in the design and training protocol of BEV perception models. We find that batch size and input resolution greatly affect performance, while lifting strategies have a more modest effect -- even a simple parameter-free lifter works well. Second, we demonstrate that radar data can provide a substantial boost to performance, helping to close the gap between camera-only and LiDAR-enabled systems. We analyze the radar usage details that lead to good performance, and invite the community to re-consider this commonly-neglected part of the sensor platform.","url_abs":"https://arxiv.org/abs/2206.07959v2","url_pdf":"https://arxiv.org/pdf/2206.07959v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-simple-baseline-for-bev-perception-without","repo_url":"https://github.com/valeoai/pointbev","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"autonomous-vehicles","task_name":"Autonomous Vehicles"},{"task_slug":"bird-s-eye-view-semantic-segmentation","task_name":"Bird's-Eye View Semantic Segmentation"},{"task_slug":"data-augmentation","task_name":"Data Augmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/bird-s-eye-view-semantic-segmentation-on-lyft","task":"Bird's-Eye View Semantic Segmentation","dataset":"Lyft Level 5","model":"Simple-BEV (EfficientNet-b4)","rank_in_archive_order":3,"of":7,"metrics":{"IoU vehicle - 224x480 - Long":"44.5","IoU vehicle - 224x480 - Short":"70.4"},"uses_additional_data":false},{"leaderboard":"/sota/bird-s-eye-view-semantic-segmentation-on-lyft","task":"Bird's-Eye View Semantic Segmentation","dataset":"Lyft Level 5","model":"Simple-BEV (ResNet-50)","rank_in_archive_order":5,"of":7,"metrics":{"IoU vehicle - 224x480 - Long":"43.6","IoU vehicle - 224x480 - Short":"70.7"},"uses_additional_data":false},{"leaderboard":"/sota/bird-s-eye-view-semantic-segmentation-on","task":"Bird's-Eye View Semantic Segmentation","dataset":"nuScenes","model":"Simple-BEV","rank_in_archive_order":3,"of":17,"metrics":{"IoU veh - 224x480 - No vis filter - 100x100 at 0.5":"36.9","IoU veh - 224x480 - Vis filter. - 100x100 at 0.5":"43.0","IoU veh - 448x800 - No vis filter - 100x100 at 0.5":"40.9","IoU veh - 448x800 - Vis filter. - 100x100 at 0.5":"46.6"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2206.07959","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}