{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multi-stage-reinforcement-learning-for-object","title":"Multi-Stage Reinforcement Learning For Object Detection","arxiv_id":"1810.10325","date":"2018-10-15","proceeding":null,"authors":["Jonas Koenig","Simon Malberg","Martin Martens","Sebastian Niehaus","Artus Krohn-Grimberghe","Arunselvan Ramaswamy"],"abstract":"We present a reinforcement learning approach for detecting objects within an\nimage. Our approach performs a step-wise deformation of a bounding box with the\ngoal of tightly framing the object. It uses a hierarchical tree-like\nrepresentation of predefined region candidates, which the agent can zoom in on.\nThis reduces the number of region candidates that must be evaluated so that the\nagent can afford to compute new feature maps before each step to enhance\ndetection quality. We compare an approach that is based purely on zoom actions\nwith one that is extended by a second refinement stage to fine-tune the\nbounding box after each zoom step. We also improve the fitting ability by\nallowing for different aspect ratios of the bounding box. Finally, we propose\ndifferent reward functions to lead to a better guidance of the agent while\nfollowing its search trajectories. Experiments indicate that each of these\nextensions leads to more correct detections. The best performing approach\ncomprises a zoom stage and a refinement stage, uses aspect-ratio modifying\nactions and is trained using a combination of three different reward metrics.","url_abs":"http://arxiv.org/abs/1810.10325v2","url_pdf":"http://arxiv.org/pdf/1810.10325v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"multi-stage-reinforcement-learning-for-object","repo_url":"https://github.com/qq456cvb/multi-stage-detection","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"object-detection-1","task_name":"object-detection"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}