{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/context-meta-reinforcement-learning-via","title":"Context Meta-Reinforcement Learning via Neuromodulation","arxiv_id":"2111.00134","date":"2021-10-30","proceeding":null,"authors":["Eseoghene Ben-Iwhiwhu","Jeffery Dick","Nicholas A. Ketz","Praveen K. Pilly","Andrea Soltoggio"],"abstract":"Meta-reinforcement learning (meta-RL) algorithms enable agents to adapt quickly to tasks from few samples in dynamic environments. Such a feat is achieved through dynamic representations in an agent's policy network (obtained via reasoning about task context, model parameter updates, or both). However, obtaining rich dynamic representations for fast adaptation beyond simple benchmark problems is challenging due to the burden placed on the policy network to accommodate different policies. This paper addresses the challenge by introducing neuromodulation as a modular component to augment a standard policy network that regulates neuronal activities in order to produce efficient dynamic representations for task adaptation. The proposed extension to the policy network is evaluated across multiple discrete and continuous control environments of increasing complexity. To prove the generality and benefits of the extension in meta-RL, the neuromodulated network was applied to two state-of-the-art meta-RL algorithms (CAVIA and PEARL). The result demonstrates that meta-RL augmented with neuromodulation produces significantly better result and richer dynamic representations in comparison to the baselines.","url_abs":"https://arxiv.org/abs/2111.00134v3","url_pdf":"https://arxiv.org/pdf/2111.00134v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"context-meta-reinforcement-learning-via","repo_url":"https://github.com/dlpbc/nm-metarl","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"context-meta-reinforcement-learning-via","repo_url":"https://github.com/soltoggio/ct-graph","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"GPL-3.0"}}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"meta-reinforcement-learning","task_name":"Meta Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"continuous-control","task_name":"continuous-control"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2111.00134","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2111.00134"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/dlpbc/nm-metarl","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/soltoggio/ct-graph","reach":{"status":"ok","spdx":"GPL-3.0"}}],"summary":{"ran_draft_wrong":2,"unverified":3},"by_repo_kind":{"official":{"samples":5,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"e937b70b07bfc9d1","entry":"deep_update_dict","repo":"dlpbc/nm-metarl","repo_kind":"official","path":"nm_oyster/launch_experiment.py","file_url":"https://github.com/dlpbc/nm-metarl/blob/HEAD/nm_oyster/launch_experiment.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e937b70b07bfc9d1"}},{"code_sha256_prefix":"b91a5961fb87cb5e","entry":"load_obj","repo":"dlpbc/nm-metarl","repo_kind":"official","path":"nm_cavia/rl/utils.py","file_url":"https://github.com/dlpbc/nm-metarl/blob/HEAD/nm_cavia/rl/utils.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b91a5961fb87cb5e"}},{"code_sha256_prefix":"3d87729bd4c217d5","entry":"get_mujoco_zip_name","repo":"dlpbc/nm-metarl","repo_kind":"official","path":"nm_oyster/docker/install_mujoco.py","file_url":"https://github.com/dlpbc/nm-metarl/blob/HEAD/nm_oyster/docker/install_mujoco.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3d87729bd4c217d5"}},{"code_sha256_prefix":"9e9b06cdc8a7e6eb","entry":"get_path_from_args","repo":"dlpbc/nm-metarl","repo_kind":"official","path":"nm_cavia/rl/utils.py","file_url":"https://github.com/dlpbc/nm-metarl/blob/HEAD/nm_cavia/rl/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9e9b06cdc8a7e6eb"}},{"code_sha256_prefix":"117528af3ffaa8d8","entry":"make_env","repo":"dlpbc/nm-metarl","repo_kind":"official","path":"nm_cavia/rl/sampler.py","file_url":"https://github.com/dlpbc/nm-metarl/blob/HEAD/nm_cavia/rl/sampler.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"117528af3ffaa8d8"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}