{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-robust-reward-machines-from-noisy","title":"Learning Robust Reward Machines from Noisy Labels","arxiv_id":"2408.14871","date":"2024-08-27","proceeding":null,"authors":["Roko Parac","Lorenzo Nodari","Leo Ardon","Daniel Furelos-Blanco","Federico Cerutti","Alessandra Russo"],"abstract":"This paper presents PROB-IRM, an approach that learns robust reward machines (RMs) for reinforcement learning (RL) agents from noisy execution traces. The key aspect of RM-driven RL is the exploitation of a finite-state machine that decomposes the agent's task into different subtasks. PROB-IRM uses a state-of-the-art inductive logic programming framework robust to noisy examples to learn RMs from noisy traces using the Bayesian posterior degree of beliefs, thus ensuring robustness against inconsistencies. Pivotal for the results is the interleaving between RM learning and policy learning: a new RM is learned whenever the RL agent generates a trace that is believed not to be accepted by the current RM. To speed up the training of the RL agent, PROB-IRM employs a probabilistic formulation of reward shaping that uses the posterior Bayesian beliefs derived from the traces. Our experimental analysis shows that PROB-IRM can learn (potentially imperfect) RMs from noisy traces and exploit them to train an RL agent to solve its tasks successfully. Despite the complexity of learning the RM from noisy traces, agents trained with PROB-IRM perform comparably to agents provided with handcrafted RMs.","url_abs":"https://arxiv.org/abs/2408.14871v1","url_pdf":"https://arxiv.org/pdf/2408.14871v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-robust-reward-machines-from-noisy","repo_url":"https://github.com/rparac/prob-irm","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"inductive-logic-programming","task_name":"Inductive logic programming"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"}],"methods":[{"method_slug":"speed","method_name":"SPEED"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2408.14871","atlas_url":"https://app.syntology.ai/?focus=2408.14871","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2408.14871"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/rparac/prob-irm","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":3,"unverified":1},"by_repo_kind":{"official":{"samples":4,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"a4bf3cfcdccde80a","entry":"get_pbs_script_base","repo":"rparac/prob-irm","repo_kind":"official","path":"submit_rcs_script.py","file_url":"https://github.com/rparac/prob-irm/blob/HEAD/submit_rcs_script.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a4bf3cfcdccde80a"}},{"code_sha256_prefix":"eaea9746cfb4f952","entry":"get_slurm_base_script","repo":"rparac/prob-irm","repo_kind":"official","path":"submit_slurm_script.py","file_url":"https://github.com/rparac/prob-irm/blob/HEAD/submit_slurm_script.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"eaea9746cfb4f952"}},{"code_sha256_prefix":"f8ac00908af2bffc","entry":"setup_rm_learner_config","repo":"rparac/prob-irm","repo_kind":"official","path":"ray_tests/hydra_RM_learning_PPO.py","file_url":"https://github.com/rparac/prob-irm/blob/HEAD/ray_tests/hydra_RM_learning_PPO.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f8ac00908af2bffc"}},{"code_sha256_prefix":"dded27282ba2609a","entry":"generate_condor_script","repo":"rparac/prob-irm","repo_kind":"official","path":"submit_condor_script.py","file_url":"https://github.com/rparac/prob-irm/blob/HEAD/submit_condor_script.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"dded27282ba2609a"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}