{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multimodal-object-detection-by-channel","title":"Multimodal Object Detection by Channel Switching and Spatial Attention","arxiv_id":null,"date":"2023-06-18","proceeding":"Conference on Computer Vision and Pattern Recognition (CVPR) 2023 6","authors":["Yue Cao","Junchi Bin","Jozsef Hamari","Erik Blasch","Zheng Liu"],"abstract":"Multimodal object detection has attracted great attention in recent years since the information specific to different modalities can complement each other and effectively improve the accuracy and stability of the detection model. However, compared to processing the inputs from a single modality, fusing information from multiple modalities can significantly increase the computational complexity of the model, thus impairing its efficiency. Therefore the multi-modal fusion module needs to be carefully designed to enhance the performance of the detection model while keeping the computational consumption low. In this paper, we propose a novel lightweight fusion module that can efficiently fuse the inputs from different modalities using channel switching and spatial attention (CSSA). The effectiveness and generalizability of the module are tested using two public multimodal datasets LLVIP and FLIR, both of which comprise paired infrared (IR) and visible (RGB) images. The experiments demonstrate that the proposed CSSA module can substantially improve the accuracy of multimodal object detection without consuming excessive computing resources.","url_abs":"https://openaccess.thecvf.com/content/CVPR2023W/PBVS/html/Cao_Multimodal_Object_Detection_by_Channel_Switching_and_Spatial_Attention_CVPRW_2023_paper.html","url_pdf":"https://openaccess.thecvf.com/content/CVPR2023W/PBVS/papers/Cao_Multimodal_Object_Detection_by_Channel_Switching_and_Spatial_Attention_CVPRW_2023_paper.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"multispectral-object-detection","task_name":"Multispectral Object Detection"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"pedestrian-detection","task_name":"Pedestrian Detection"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/multispectral-object-detection-on-flir-1","task":"Multispectral Object Detection","dataset":"FLIR","model":"CSSA","rank_in_archive_order":8,"of":18,"metrics":{"mAP":"41.3%","mAP50":"79.2%"},"uses_additional_data":false},{"leaderboard":"/sota/multispectral-object-detection-on-flir-1","task":"Multispectral Object Detection","dataset":"FLIR","model":"ProbEn","rank_in_archive_order":10,"of":18,"metrics":{"mAP":"37.9%","mAP50":"75.5%"},"uses_additional_data":false},{"leaderboard":"/sota/multispectral-object-detection-on-flir-1","task":"Multispectral Object Detection","dataset":"FLIR","model":"GAFF","rank_in_archive_order":11,"of":18,"metrics":{"mAP":"37.4%","mAP50":"74.6%"},"uses_additional_data":false},{"leaderboard":"/sota/multispectral-object-detection-on-flir-1","task":"Multispectral Object Detection","dataset":"FLIR","model":"Halfway Fusion","rank_in_archive_order":18,"of":18,"metrics":{"mAP":"35.8%"},"uses_additional_data":false},{"leaderboard":"/sota/pedestrian-detection-on-llvip","task":"Pedestrian Detection","dataset":"LLVIP","model":"CSSA","rank_in_archive_order":8,"of":15,"metrics":{"AP":"0.592"},"uses_additional_data":false},{"leaderboard":"/sota/pedestrian-detection-on-llvip","task":"Pedestrian Detection","dataset":"LLVIP","model":"GAFF","rank_in_archive_order":9,"of":15,"metrics":{"AP":"0.558"},"uses_additional_data":false},{"leaderboard":"/sota/pedestrian-detection-on-llvip","task":"Pedestrian Detection","dataset":"LLVIP","model":"Halfway Fusion","rank_in_archive_order":10,"of":15,"metrics":{"AP":"0.551"},"uses_additional_data":false},{"leaderboard":"/sota/pedestrian-detection-on-llvip","task":"Pedestrian Detection","dataset":"LLVIP","model":"ProbEn","rank_in_archive_order":12,"of":15,"metrics":{"AP":"0.515"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}