{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/interactive-learning-with-corrective-feedback","title":"Interactive Learning with Corrective Feedback for Policies based on Deep Neural Networks","arxiv_id":"1810.00466","date":"2018-09-30","proceeding":null,"authors":["Rodrigo Pérez-Dattari","Carlos Celemin","Javier Ruiz-del-Solar","Jens Kober"],"abstract":"Deep Reinforcement Learning (DRL) has become a powerful strategy to solve\ncomplex decision making problems based on Deep Neural Networks (DNNs). However,\nit is highly data demanding, so unfeasible in physical systems for most\napplications. In this work, we approach an alternative Interactive Machine\nLearning (IML) strategy for training DNN policies based on human corrective\nfeedback, with a method called Deep COACH (D-COACH). This approach not only\ntakes advantage of the knowledge and insights of human teachers as well as the\npower of DNNs, but also has no need of a reward function (which sometimes\nimplies the need of external perception for computing rewards). We combine Deep\nLearning with the COrrective Advice Communicated by Humans (COACH) framework,\nin which non-expert humans shape policies by correcting the agent's actions\nduring execution. The D-COACH framework has the potential to solve complex\nproblems without much data or time required. Experimental results validated the\nefficiency of the framework in three different problems (two simulated, one\nwith a real robot), with state spaces of low and high dimensions, showing the\ncapacity to successfully learn policies for continuous action spaces like in\nthe Car Racing and Cart-Pole problems faster than with DRL.","url_abs":"http://arxiv.org/abs/1810.00466v1","url_pdf":"http://arxiv.org/pdf/1810.00466v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"interactive-learning-with-corrective-feedback","repo_url":"https://github.com/rperezdattari/Interactive-Learning-with-Corrective-Feedback-for-Policies-based-on-Deep-Neural-Networks","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"car-racing","task_name":"Car Racing"},{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}