{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/fused-dnn-a-deep-neural-network-fusion","title":"Fused DNN: A deep neural network fusion approach to fast and robust pedestrian detection","arxiv_id":"1610.03466","date":"2016-10-11","proceeding":null,"authors":["Xianzhi Du","Mostafa El-Khamy","Jungwon Lee","Larry S. Davis"],"abstract":"We propose a deep neural network fusion architecture for fast and robust\npedestrian detection. The proposed network fusion architecture allows for\nparallel processing of multiple networks for speed. A single shot deep\nconvolutional network is trained as a object detector to generate all possible\npedestrian candidates of different sizes and occlusions. This network outputs a\nlarge variety of pedestrian candidates to cover the majority of ground-truth\npedestrians while also introducing a large number of false positives. Next,\nmultiple deep neural networks are used in parallel for further refinement of\nthese pedestrian candidates. We introduce a soft-rejection based network fusion\nmethod to fuse the soft metrics from all networks together to generate the\nfinal confidence scores. Our method performs better than existing\nstate-of-the-arts, especially when detecting small-size and occluded\npedestrians. Furthermore, we propose a method for integrating pixel-wise\nsemantic segmentation network into the network fusion architecture as a\nreinforcement to the pedestrian detector. The approach outperforms\nstate-of-the-art methods on most protocols on Caltech Pedestrian dataset, with\nsignificant boosts on several protocols. It is also faster than all other\nmethods.","url_abs":"http://arxiv.org/abs/1610.03466v2","url_pdf":"http://arxiv.org/pdf/1610.03466v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"pedestrian-detection","task_name":"Pedestrian Detection"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/pedestrian-detection-on-caltech","task":"Pedestrian Detection","dataset":"Caltech","model":"F-DNN+SS","rank_in_archive_order":21,"of":33,"metrics":{"Reasonable Miss Rate":"8.18"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}