{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/imitation-learning-from-purified","title":"Imitation Learning from Purified Demonstrations","arxiv_id":"2310.07143","date":"2023-10-11","proceeding":null,"authors":["Yunke Wang","Minjing Dong","Yukun Zhao","Bo Du","Chang Xu"],"abstract":"Imitation learning has emerged as a promising approach for addressing sequential decision-making problems, with the assumption that expert demonstrations are optimal. However, in real-world scenarios, most demonstrations are often imperfect, leading to challenges in the effectiveness of imitation learning. While existing research has focused on optimizing with imperfect demonstrations, the training typically requires a certain proportion of optimal demonstrations to guarantee performance. To tackle these problems, we propose to purify the potential noises in imperfect demonstrations first, and subsequently conduct imitation learning from these purified demonstrations. Motivated by the success of diffusion model, we introduce a two-step purification via diffusion process. In the first step, we apply a forward diffusion process to smooth potential noises in imperfect demonstrations by introducing additional noise. Subsequently, a reverse generative process is utilized to recover the optimal demonstration from the diffused ones. We provide theoretical evidence supporting our approach, demonstrating that the distance between the purified and optimal demonstration can be bounded. Empirical results on MuJoCo and RoboSuite demonstrate the effectiveness of our method from different aspects.","url_abs":"https://arxiv.org/abs/2310.07143v2","url_pdf":"https://arxiv.org/pdf/2310.07143v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"imitation-learning-from-purified","repo_url":"https://github.com/yunke-wang/dp-il","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"imitation-learning","task_name":"Imitation Learning"},{"task_slug":"mujoco","task_name":"MuJoCo"},{"task_slug":"sequential-decision-making","task_name":"Sequential Decision Making"}],"methods":[{"method_slug":"diffusion","method_name":"Diffusion"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2310.07143","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2310.07143"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/yunke-wang/dp-il","reach":null},{"provenance":"deterministic:regex_extraction","url":"https://github.com/abarankab/DDPM","reach":null}],"summary":{"ran_honours":2,"ran":2,"unverified":1},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1},"found_in_text":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"b67dc049ab8bd6af","entry":"generate_linear_schedule","repo":"yunke-wang/dp-il","repo_kind":"official","path":"code/core/ddpm.py","file_url":"https://github.com/yunke-wang/dp-il/blob/HEAD/code/core/ddpm.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":2,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"b67dc049ab8bd6af"}},{"code_sha256_prefix":"9b9406256db3fe0c","entry":"EMA","repo":"abarankab/DDPM","repo_kind":"found_in_text","path":"ddpm/diffusion.py","file_url":"https://github.com/abarankab/DDPM/blob/HEAD/ddpm/diffusion.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9b9406256db3fe0c"}},{"code_sha256_prefix":"194e1130f6ac20c5","entry":"GaussianDiffusion","repo":"abarankab/DDPM","repo_kind":"found_in_text","path":"ddpm/diffusion.py","file_url":"https://github.com/abarankab/DDPM/blob/HEAD/ddpm/diffusion.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"194e1130f6ac20c5"}},{"code_sha256_prefix":"614a03f150609f46","entry":"generate_cosine_schedule","repo":"yunke-wang/dp-il","repo_kind":"official","path":"code/core/ddpm.py","file_url":"https://github.com/yunke-wang/dp-il/blob/HEAD/code/core/ddpm.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"614a03f150609f46"}},{"code_sha256_prefix":"09c8479d9a5b3e06","entry":"extract","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"09c8479d9a5b3e06"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}