{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/rcfusion-fusing-4-d-radar-and-camera-with","title":"RCFusion: Fusing 4-D Radar and Camera With Bird’s-Eye View Features for 3-D Object Detection","arxiv_id":null,"date":"2023-05-23","proceeding":"IEEE Transactions on Instrumentation and Measurement 2023 5","authors":["Lianqing Zheng","Sen Li","Bin Tan","Long Yan","Sihan Chen","Libo Huang","Jie Bai","Xichan Zhu","Zhixiong Ma"],"abstract":"Camera and millimeter-wave (MMW) radar fusion is essential for accurate and robust autonomous driving systems. With the advancement of radar technology, next-generation high-resolution automotive radar, i.e., 4-D radar, has emerged. In addition to the target range, azimuth, and Doppler velocity measurements of traditional radar, 4-D radar provides elevation measurement to create a denser “point cloud.” In this study, we propose a camera and 4-D radar fusion network called RCFusion, which achieves multimodal feature fusion under a unified bird’s-eye view (BEV) space to accomplish 3-D object detection tasks. In the camera stream, multiscale feature maps are obtained by the image backbone and feature pyramid network (FPN); they are then converted into orthographic feature maps by an orthographic feature transform (OFT). Next, enhanced and fine-grained image BEV features are obtained via a designed shared attention encoder. Meanwhile, in the 4-D radar stream, a newly designed component named radar PillarNet efficiently encodes the radar features to generate radar pseudo-images, which are fed into the point cloud backbone to create radar BEV features. An interactive attention module (IAM) is proposed for the fusion stage, which outputs a valid fusion of the two-modal BEV features. Finally, a generic detection head predicts the object classes and locations. The proposed RCFusion is validated on the TJ4DRadSet and view-of-delft (VoD) datasets. The experimental results and analysis show that the proposed method can effectively fuse camera and 4-D radar features to achieve robust detection performance.","url_abs":"https://ieeexplore.ieee.org/document/10138035","url_pdf":"https://ieeexplore.ieee.org/document/10138035","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"3d-object-detection","task_name":"3D Object Detection"},{"task_slug":"3d-object-detection-roi","task_name":"3D Object Detection (RoI)"},{"task_slug":"autonomous-driving","task_name":"Autonomous Driving"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/3d-object-detection-on-view-of-delft-val","task":"3D Object Detection","dataset":"View-of-Delft (val)","model":"RCFusion","rank_in_archive_order":7,"of":11,"metrics":{"mAP":"49.7"},"uses_additional_data":false},{"leaderboard":"/sota/3d-object-detection-on-view-of-delft-val","task":"3D Object Detection","dataset":"View-of-Delft (val)","model":"RadarPillarNet","rank_in_archive_order":10,"of":11,"metrics":{"mAP":"46.0"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}