{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-from-imperfect-demonstrations-from","title":"Learning from Imperfect Demonstrations from Agents with Varying Dynamics","arxiv_id":"2103.05910","date":"2021-03-10","proceeding":null,"authors":["Zhangjie Cao","Dorsa Sadigh"],"abstract":"Imitation learning enables robots to learn from demonstrations. Previous imitation learning algorithms usually assume access to optimal expert demonstrations. However, in many real-world applications, this assumption is limiting. Most collected demonstrations are not optimal or are produced by an agent with slightly different dynamics. We therefore address the problem of imitation learning when the demonstrations can be sub-optimal or be drawn from agents with varying dynamics. We develop a metric composed of a feasibility score and an optimality score to measure how useful a demonstration is for imitation learning. The proposed score enables learning from more informative demonstrations, and disregarding the less relevant demonstrations. Our experiments on four environments in simulation and on a real robot show improved learned policies with higher expected return.","url_abs":"https://arxiv.org/abs/2103.05910v1","url_pdf":"https://arxiv.org/pdf/2103.05910v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-from-imperfect-demonstrations-from","repo_url":"https://github.com/Stanford-ILIAD/Learn-Imperfect-Varying-Dynamics","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"imitation-learning","task_name":"Imitation Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2103.05910","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2103.05910"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Stanford-ILIAD/Learn-Imperfect-Varying-Dynamics","reach":null}],"summary":{"ran_fixture":1,"unverified":2},"by_repo_kind":{"listed":{"samples":3,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"918891131c62a121","entry":"conjugate_gradients","repo":"Stanford-ILIAD/Learn-Imperfect-Varying-Dynamics","repo_kind":"listed","path":"trpo.py","file_url":"https://github.com/Stanford-ILIAD/Learn-Imperfect-Varying-Dynamics/blob/HEAD/trpo.py","link_basis":"first_harvest_node","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"918891131c62a121"}},{"code_sha256_prefix":"977dcd409cef091b","entry":"linesearch","repo":"Stanford-ILIAD/Learn-Imperfect-Varying-Dynamics","repo_kind":"listed","path":"trpo.py","file_url":"https://github.com/Stanford-ILIAD/Learn-Imperfect-Varying-Dynamics/blob/HEAD/trpo.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"977dcd409cef091b"}},{"code_sha256_prefix":"9baac4ef495e3380","entry":"trpo_step","repo":"Stanford-ILIAD/Learn-Imperfect-Varying-Dynamics","repo_kind":"listed","path":"trpo.py","file_url":"https://github.com/Stanford-ILIAD/Learn-Imperfect-Varying-Dynamics/blob/HEAD/trpo.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9baac4ef495e3380"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}