{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multispectral-pedestrian-detection-via","title":"Multispectral Pedestrian Detection via Simultaneous Detection and Segmentation","arxiv_id":"1808.04818","date":"2018-08-14","proceeding":null,"authors":["Chengyang Li","Dan Song","Ruofeng Tong","Min Tang"],"abstract":"Multispectral pedestrian detection has attracted increasing attention from\nthe research community due to its crucial competence for many around-the-clock\napplications (e.g., video surveillance and autonomous driving), especially\nunder insufficient illumination conditions. We create a human baseline over the\nKAIST dataset and reveal that there is still a large gap between current top\ndetectors and human performance. To narrow this gap, we propose a network\nfusion architecture, which consists of a multispectral proposal network to\ngenerate pedestrian proposals, and a subsequent multispectral classification\nnetwork to distinguish pedestrian instances from hard negatives. The unified\nnetwork is learned by jointly optimizing pedestrian detection and semantic\nsegmentation tasks. The final detections are obtained by integrating the\noutputs from different modalities as well as the two stages. The approach\nsignificantly outperforms state-of-the-art methods on the KAIST dataset while\nremain fast. Additionally, we contribute a sanitized version of training\nannotations for the KAIST dataset, and examine the effects caused by different\nkinds of annotation errors. Future research of this problem will benefit from\nthe sanitized version which eliminates the interference of annotation errors.","url_abs":"http://arxiv.org/abs/1808.04818v1","url_pdf":"http://arxiv.org/pdf/1808.04818v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multispectral-pedestrian-detection-via","repo_url":"https://github.com/Li-Chengyang/MSDS-RCNN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"autonomous-driving","task_name":"Autonomous Driving"},{"task_slug":"multispectral-object-detection","task_name":"Multispectral Object Detection"},{"task_slug":"pedestrian-detection","task_name":"Pedestrian Detection"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/multispectral-object-detection-on-kaist","task":"Multispectral Object Detection","dataset":"KAIST Multispectral Pedestrian Detection Benchmark","model":"MSDS-R-CNN","rank_in_archive_order":9,"of":17,"metrics":{"All Miss Rate":"34.15"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1808.04818","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}