{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/scale-aware-fast-r-cnn-for-pedestrian","title":"Scale-aware Fast R-CNN for Pedestrian Detection","arxiv_id":"1510.08160","date":"2015-10-28","proceeding":null,"authors":["Jianan Li","Xiaodan Liang","ShengMei Shen","Tingfa Xu","Jiashi Feng","Shuicheng Yan"],"abstract":"In this work, we consider the problem of pedestrian detection in natural\nscenes. Intuitively, instances of pedestrians with different spatial scales may\nexhibit dramatically different features. Thus, large variance in instance\nscales, which results in undesirable large intra-category variance in features,\nmay severely hurt the performance of modern object instance detection methods.\nWe argue that this issue can be substantially alleviated by the\ndivide-and-conquer philosophy. Taking pedestrian detection as an example, we\nillustrate how we can leverage this philosophy to develop a Scale-Aware Fast\nR-CNN (SAF R-CNN) framework. The model introduces multiple built-in\nsub-networks which detect pedestrians with scales from disjoint ranges. Outputs\nfrom all the sub-networks are then adaptively combined to generate the final\ndetection results that are shown to be robust to large variance in instance\nscales, via a gate function defined over the sizes of object proposals.\nExtensive evaluations on several challenging pedestrian detection datasets well\ndemonstrate the effectiveness of the proposed SAF R-CNN. Particularly, our\nmethod achieves state-of-the-art performance on Caltech, INRIA, and ETH, and\nobtains competitive results on KITTI.","url_abs":"http://arxiv.org/abs/1510.08160v3","url_pdf":"http://arxiv.org/pdf/1510.08160v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"pedestrian-detection","task_name":"Pedestrian Detection"},{"task_slug":"philosophy","task_name":"Philosophy"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/pedestrian-detection-on-caltech","task":"Pedestrian Detection","dataset":"Caltech","model":"SA-FastRCNN","rank_in_archive_order":23,"of":33,"metrics":{"Reasonable Miss Rate":"9.68"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1510.08160","atlas_url":"https://app.syntology.ai/?focus=1510.08160","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}