{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/accurately-and-efficiently-interpreting-human","title":"Accurately and Efficiently Interpreting Human-Robot Instructions of Varying Granularities","arxiv_id":"1704.06616","date":"2017-04-21","proceeding":null,"authors":["Dilip Arumugam","Siddharth Karamcheti","Nakul Gopalan","Lawson L. S. Wong","Stefanie Tellex"],"abstract":"Humans can ground natural language commands to tasks at both abstract and\nfine-grained levels of specificity. For instance, a human forklift operator can\nbe instructed to perform a high-level action, like \"grab a pallet\" or a\nlow-level action like \"tilt back a little bit.\" While robots are also capable\nof grounding language commands to tasks, previous methods implicitly assume\nthat all commands and tasks reside at a single, fixed level of abstraction.\nAdditionally, methods that do not use multiple levels of abstraction encounter\ninefficient planning and execution times as they solve tasks at a single level\nof abstraction with large, intractable state-action spaces closely resembling\nreal world complexity. In this work, by grounding commands to all the tasks or\nsubtasks available in a hierarchical planning framework, we arrive at a model\ncapable of interpreting language at multiple levels of specificity ranging from\ncoarse to more granular. We show that the accuracy of the grounding procedure\nis improved when simultaneously inferring the degree of abstraction in language\nused to communicate the task. Leveraging hierarchy also improves efficiency:\nour proposed approach enables a robot to respond to a command within one second\non 90% of our tasks, while baselines take over twenty seconds on half the\ntasks. Finally, we demonstrate that a real, physical robot can ground commands\nat multiple levels of abstraction allowing it to efficiently plan different\nsubtasks within the same planning hierarchy.","url_abs":"http://arxiv.org/abs/1704.06616v2","url_pdf":"http://arxiv.org/pdf/1704.06616v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"accurately-and-efficiently-interpreting-human","repo_url":"https://github.com/h2r/GLAMDP","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"specificity","task_name":"Specificity"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1704.06616","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}