{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/caff-dino-multi-spectral-object-detection","title":"CAFF-DINO: Multi-spectral object detection transformers with cross-attention features fusion","arxiv_id":null,"date":"2024-09-27","proceeding":"IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) 2024 9","authors":["Kevin Helvig","Baptiste Abeloos","Pauline Trouve-Peloux"],"abstract":"Object detection on images can find benefit from coupling multiple spectra, each presenting specific useful features. However, building an efficient architecture coupling the different modalities is a complex task. Transformers, due to their ability to extract meaningful correlations between the different regions of the inputs appear as a promising way to perform features fusion across different spectra. This work presents a multi-spectral object detection architecture based on cross-attention features fusion (CAFF), combined with a transformer based detector (DINO). We demonstrate here the performance of the proposed approach in object detection compared with state-of-the-art approaches, on infrared-visible multi-spectral datasets. Moreover the robustness to systematic misalignment between image pairs is studied. The proposed approach is generic to any mono-spectrum transformer based detectors. The model developed in this study will be available in a dedicated github repository.","url_abs":"https://ieeexplore.ieee.org/stamp/stamp.jsp?arnumber=10678534","url_pdf":"https://openaccess.thecvf.com/content/CVPR2024W/PBVS/papers/Helvig_CAFF-DINO_Multi-spectral_Object_Detection_Transformers_with_Cross-attention_Features_Fusion_CVPRW_2024_paper.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"multispectral-object-detection","task_name":"Multispectral Object Detection"},{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"pedestrian-detection","task_name":"Pedestrian Detection"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/multispectral-object-detection-on-flir-1","task":"Multispectral Object Detection","dataset":"FLIR","model":"CAFF-DINO","rank_in_archive_order":3,"of":18,"metrics":{"mAP":"50.5%","mAP50":"85.5%"},"uses_additional_data":false},{"leaderboard":"/sota/pedestrian-detection-on-llvip","task":"Pedestrian Detection","dataset":"LLVIP","model":"CAFF-DINO","rank_in_archive_order":2,"of":15,"metrics":{"AP":"0.685"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}