{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/context-aware-cnns-for-person-head-detection","title":"Context-aware CNNs for person head detection","arxiv_id":"1511.07917","date":"2015-11-24","proceeding":"ICCV 2015 12","authors":["Tuan-Hung Vu","Anton Osokin","Ivan Laptev"],"abstract":"Person detection is a key problem for many computer vision tasks. While face\ndetection has reached maturity, detecting people under a full variation of\ncamera view-points, human poses, lighting conditions and occlusions is still a\ndifficult challenge. In this work we focus on detecting human heads in natural\nscenes. Starting from the recent local R-CNN object detector, we extend it with\ntwo types of contextual cues. First, we leverage person-scene relations and\npropose a Global CNN model trained to predict positions and scales of heads\ndirectly from the full image. Second, we explicitly model pairwise relations\namong objects and train a Pairwise CNN model using a structured-output\nsurrogate loss. The Local, Global and Pairwise models are combined into a joint\nCNN framework. To train and test our full model, we introduce a large dataset\ncomposed of 369,846 human heads annotated in 224,740 movie frames. We evaluate\nour method and demonstrate improvements of person head detection against\nseveral recent baselines in three datasets. We also show improvements of the\ndetection speed provided by our model.","url_abs":"http://arxiv.org/abs/1511.07917v1","url_pdf":"http://arxiv.org/pdf/1511.07917v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"context-aware-cnns-for-person-head-detection","repo_url":"https://github.com/aosokin/cnn_head_detection","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"torch","reach":null}],"tasks":[{"task_slug":"face-detection","task_name":"Face Detection"},{"task_slug":"head-detection","task_name":"Head Detection"},{"task_slug":"human-detection","task_name":"Human Detection"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"r-cnn","method_name":"R-CNN"},{"method_slug":"speed","method_name":"SPEED"},{"method_slug":"svm","method_name":"SVM"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1511.07917","atlas_url":"https://app.syntology.ai/?focus=1511.07917","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}