{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-cross-modal-deep-representations-for","title":"Learning Cross-Modal Deep Representations for Robust Pedestrian Detection","arxiv_id":"1704.02431","date":"2017-04-08","proceeding":"CVPR 2017 7","authors":["Dan Xu","Wanli Ouyang","Elisa Ricci","Xiaogang Wang","Nicu Sebe"],"abstract":"This paper presents a novel method for detecting pedestrians under adverse\nillumination conditions. Our approach relies on a novel cross-modality learning\nframework and it is based on two main phases. First, given a multimodal\ndataset, a deep convolutional network is employed to learn a non-linear\nmapping, modeling the relations between RGB and thermal data. Then, the learned\nfeature representations are transferred to a second deep network, which\nreceives as input an RGB image and outputs the detection results. In this way,\nfeatures which are both discriminative and robust to bad illumination\nconditions are learned. Importantly, at test time, only the second pipeline is\nconsidered and no thermal data are required. Our extensive evaluation\ndemonstrates that the proposed approach outperforms the state-of- the-art on\nthe challenging KAIST multispectral pedestrian dataset and it is competitive\nwith previous methods on the popular Caltech dataset.","url_abs":"http://arxiv.org/abs/1704.02431v2","url_pdf":"http://arxiv.org/pdf/1704.02431v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-cross-modal-deep-representations-for","repo_url":"https://github.com/SoonminHwang/rgbt-ped-detection","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}},{"paper_slug":"learning-cross-modal-deep-representations-for","repo_url":"https://github.com/danxuhk/CMT-CNN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"pedestrian-detection","task_name":"Pedestrian Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1704.02431","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}