{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/hierarchical-object-detection-with-deep","title":"Hierarchical Object Detection with Deep Reinforcement Learning","arxiv_id":"1611.03718","date":"2016-11-11","proceeding":null,"authors":["Miriam Bellver","Xavier Giro-i-Nieto","Ferran Marques","Jordi Torres"],"abstract":"We present a method for performing hierarchical object detection in images\nguided by a deep reinforcement learning agent. The key idea is to focus on\nthose parts of the image that contain richer information and zoom on them. We\ntrain an intelligent agent that, given an image window, is capable of deciding\nwhere to focus the attention among five different predefined region candidates\n(smaller windows). This procedure is iterated providing a hierarchical image\nanalysis.We compare two different candidate proposal strategies to guide the\nobject search: with and without overlap. Moreover, our work compares two\ndifferent strategies to extract features from a convolutional neural network\nfor each region proposal: a first one that computes new feature maps for each\nregion proposal, and a second one that computes the feature maps for the whole\nimage to later generate crops for each region proposal. Experiments indicate\nbetter results for the overlapping candidate proposal strategy and a loss of\nperformance for the cropped image features due to the loss of spatial\nresolution. We argue that, while this loss seems unavoidable when working with\nlarge amounts of object candidates, the much more reduced amount of region\nproposals generated by our reinforcement learning agent allows considering to\nextract features for each location without sharing convolutional computation\namong regions.","url_abs":"http://arxiv.org/abs/1611.03718v2","url_pdf":"http://arxiv.org/pdf/1611.03718v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"hierarchical-object-detection-with-deep","repo_url":"https://github.com/imatge-upc/detection-2016-nipsws","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"region-proposal","task_name":"Region Proposal"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"object-detection-1","task_name":"object-detection"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1611.03718","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}