{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-dataset-for-developing-and-benchmarking","title":"A Dataset for Developing and Benchmarking Active Vision","arxiv_id":"1702.08272","date":"2017-02-27","proceeding":null,"authors":["Phil Ammirato","Patrick Poirson","Eunbyung Park","Jana Kosecka","Alexander C. Berg"],"abstract":"We present a new public dataset with a focus on simulating robotic vision\ntasks in everyday indoor environments using real imagery. The dataset includes\n20,000+ RGB-D images and 50,000+ 2D bounding boxes of object instances densely\ncaptured in 9 unique scenes. We train a fast object category detector for\ninstance detection on our data. Using the dataset we show that, although\nincreasingly accurate and fast, the state of the art for object detection is\nstill severely impacted by object scale, occlusion, and viewing direction all\nof which matter for robotics applications. We next validate the dataset for\nsimulating active vision, and use the dataset to develop and evaluate a\ndeep-network-based system for next best move prediction for object\nclassification using reinforcement learning. Our dataset is available for\ndownload at cs.unc.edu/~ammirato/active_vision_dataset_website/.","url_abs":"http://arxiv.org/abs/1702.08272v2","url_pdf":"http://arxiv.org/pdf/1702.08272v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"benchmarking","task_name":"Benchmarking"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"object-detection-1","task_name":"object-detection"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[{"slug":"avd","name":"AVD","full_name":"Active Vision Dataset"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1702.08272","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}