{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-visual-servoing-with-deep-features","title":"Learning Visual Servoing with Deep Features and Fitted Q-Iteration","arxiv_id":"1703.11000","date":"2017-03-31","proceeding":null,"authors":["Alex X. Lee","Sergey Levine","Pieter Abbeel"],"abstract":"Visual servoing involves choosing actions that move a robot in response to\nobservations from a camera, in order to reach a goal configuration in the\nworld. Standard visual servoing approaches typically rely on manually designed\nfeatures and analytical dynamics models, which limits their generalization\ncapability and often requires extensive application-specific feature and model\nengineering. In this work, we study how learned visual features, learned\npredictive dynamics models, and reinforcement learning can be combined to learn\nvisual servoing mechanisms. We focus on target following, with the goal of\ndesigning algorithms that can learn a visual servo using low amounts of data of\nthe target in question, to enable quick adaptation to new targets. Our approach\nis based on servoing the camera in the space of learned visual features, rather\nthan image pixels or manually-designed keypoints. We demonstrate that standard\ndeep features, in our case taken from a model trained for object\nclassification, can be used together with a bilinear predictive model to learn\nan effective visual servo that is robust to visual variation, changes in\nviewing angle and appearance, and occlusions. A key component of our approach\nis to use a sample-efficient fitted Q-iteration algorithm to learn which\nfeatures are best suited for the task at hand. We show that we can learn an\neffective visual servo on a complex synthetic car following benchmark using\njust 20 training trajectory samples for reinforcement learning. We demonstrate\nsubstantial improvement over a conventional approach based on image pixels or\nhand-designed keypoints, and we show an improvement in sample-efficiency of\nmore than two orders of magnitude over standard model-free deep reinforcement\nlearning algorithms. Videos are available at\nhttp://rll.berkeley.edu/visual_servoing .","url_abs":"http://arxiv.org/abs/1703.11000v2","url_pdf":"http://arxiv.org/pdf/1703.11000v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-visual-servoing-with-deep-features","repo_url":"https://github.com/alexlee-gk/citysim3d","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"learning-visual-servoing-with-deep-features","repo_url":"https://github.com/alexlee-gk/visual_dynamics","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}