{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/pedestrian-detection-aided-by-deep-learning","title":"Pedestrian Detection aided by Deep Learning Semantic Tasks","arxiv_id":"1412.0069","date":"2014-11-29","proceeding":"CVPR 2015 6","authors":["Yonglong Tian","Ping Luo","Xiaogang Wang","Xiaoou Tang"],"abstract":"Deep learning methods have achieved great success in pedestrian detection,\nowing to its ability to learn features from raw pixels. However, they mainly\ncapture middle-level representations, such as pose of pedestrian, but confuse\npositive with hard negative samples, which have large ambiguity, e.g. the shape\nand appearance of `tree trunk' or `wire pole' are similar to pedestrian in\ncertain viewpoint. This ambiguity can be distinguished by high-level\nrepresentation. To this end, this work jointly optimizes pedestrian detection\nwith semantic tasks, including pedestrian attributes (e.g. `carrying backpack')\nand scene attributes (e.g. `road', `tree', and `horizontal'). Rather than\nexpensively annotating scene attributes, we transfer attributes information\nfrom existing scene segmentation datasets to the pedestrian dataset, by\nproposing a novel deep model to learn high-level features from multiple tasks\nand multiple data sources. Since distinct tasks have distinct convergence rates\nand data from different datasets have different distributions, a multi-task\nobjective function is carefully designed to coordinate tasks and reduce\ndiscrepancies among datasets. The importance coefficients of tasks and network\nparameters in this objective function can be iteratively estimated. Extensive\nevaluations show that the proposed approach outperforms the state-of-the-art on\nthe challenging Caltech and ETH datasets, where it reduces the miss rates of\nprevious deep models by 17 and 5.5 percent, respectively.","url_abs":"http://arxiv.org/abs/1412.0069v1","url_pdf":"http://arxiv.org/pdf/1412.0069v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"deep-learning","task_name":"Deep Learning"},{"task_slug":"pedestrian-detection","task_name":"Pedestrian Detection"},{"task_slug":"scene-segmentation","task_name":"Scene Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/pedestrian-detection-on-caltech","task":"Pedestrian Detection","dataset":"Caltech","model":"TA-CNN","rank_in_archive_order":30,"of":33,"metrics":{"Reasonable Miss Rate":"20.9"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1412.0069","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}