{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/one-shot-learning-of-multi-step-tasks-from","title":"One-Shot Learning of Multi-Step Tasks from Observation via Activity Localization in Auxiliary Video","arxiv_id":"1806.11244","date":"2018-06-29","proceeding":null,"authors":["Wonjoon Goo","Scott Niekum"],"abstract":"Due to burdensome data requirements, learning from demonstration often falls\nshort of its promise to allow users to quickly and naturally program robots.\nDemonstrations are inherently ambiguous and incomplete, making correct\ngeneralization to unseen situations difficult without a large number of\ndemonstrations in varying conditions. By contrast, humans are often able to\nlearn complex tasks from a single demonstration (typically observations without\naction labels) by leveraging context learned over a lifetime. Inspired by this\ncapability, our goal is to enable robots to perform one-shot learning of\nmulti-step tasks from observation by leveraging auxiliary video data as\ncontext. Our primary contribution is a novel system that achieves this goal by:\n(1) using a single user-segmented demonstration to define the primitive actions\nthat comprise a task, (2) localizing additional examples of these actions in\nunsegmented auxiliary videos via a metalearning-based approach, (3) using these\nadditional examples to learn a reward function for each action, and (4)\nperforming reinforcement learning on top of the inferred reward functions to\nlearn action policies that can be combined to accomplish the task. We\nempirically demonstrate that a robot can learn multi-step tasks more\neffectively when provided auxiliary video, and that performance greatly\nimproves when localizing individual actions, compared to learning from\nunsegmented videos.","url_abs":"http://arxiv.org/abs/1806.11244v3","url_pdf":"http://arxiv.org/pdf/1806.11244v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"one-shot-learning-of-multi-step-tasks-from","repo_url":"https://github.com/hiwonjoon/ICRA2019-Activity-Localize","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"one-shot-learning","task_name":"One-Shot Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"task-2","task_name":"Task 2"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1806.11244","atlas_url":"https://app.syntology.ai/?focus=1806.11244","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}