{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/align-rudder-learning-from-few-demonstrations","title":"Align-RUDDER: Learning From Few Demonstrations by Reward Redistribution","arxiv_id":"2009.14108","date":"2020-09-29","proceeding":null,"authors":["Vihang P. Patil","Markus Hofmarcher","Marius-Constantin Dinu","Matthias Dorfer","Patrick M. Blies","Johannes Brandstetter","Jose A. Arjona-Medina","Sepp Hochreiter"],"abstract":"Reinforcement learning algorithms require many samples when solving complex hierarchical tasks with sparse and delayed rewards. For such complex tasks, the recently proposed RUDDER uses reward redistribution to leverage steps in the Q-function that are associated with accomplishing sub-tasks. However, often only few episodes with high rewards are available as demonstrations since current exploration strategies cannot discover them in reasonable time. In this work, we introduce Align-RUDDER, which utilizes a profile model for reward redistribution that is obtained from multiple sequence alignment of demonstrations. Consequently, Align-RUDDER employs reward redistribution effectively and, thereby, drastically improves learning on few demonstrations. Align-RUDDER outperforms competitors on complex artificial tasks with delayed rewards and few demonstrations. On the Minecraft ObtainDiamond task, Align-RUDDER is able to mine a diamond, though not frequently. Code is available at https://github.com/ml-jku/align-rudder. YouTube: https://youtu.be/HO-_8ZUl-UY","url_abs":"https://arxiv.org/abs/2009.14108v2","url_pdf":"https://arxiv.org/pdf/2009.14108v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"align-rudder-learning-from-few-demonstrations","repo_url":"https://github.com/ml-jku/align-rudder","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"general-reinforcement-learning","task_name":"General Reinforcement Learning"},{"task_slug":"minecraft","task_name":"Minecraft"},{"task_slug":"multiple-sequence-alignment","task_name":"Multiple Sequence Alignment"},{"task_slug":"safe-exploration","task_name":"Safe Exploration"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2009.14108","atlas_url":"https://app.syntology.ai/?focus=2009.14108","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2009.14108"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ml-jku/align-rudder","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":3},"by_repo_kind":{"official":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"4385bd0521d18b6d","entry":"create_fasta_sequences","repo":"ml-jku/align-rudder","repo_kind":"official","path":"align_rudder/alignment/align.py","file_url":"https://github.com/ml-jku/align-rudder/blob/HEAD/align_rudder/alignment/align.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4385bd0521d18b6d"}},{"code_sha256_prefix":"e419a78fa5fd2c84","entry":"create_scoring_matrix","repo":"ml-jku/align-rudder","repo_kind":"official","path":"align_rudder/alignment/align.py","file_url":"https://github.com/ml-jku/align-rudder/blob/HEAD/align_rudder/alignment/align.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e419a78fa5fd2c84"}},{"code_sha256_prefix":"c8b725ba492a3f9c","entry":"get_alignment","repo":"ml-jku/align-rudder","repo_kind":"official","path":"align_rudder/alignment/align.py","file_url":"https://github.com/ml-jku/align-rudder/blob/HEAD/align_rudder/alignment/align.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c8b725ba492a3f9c"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}