{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/improving-reinforcement-learning-training","title":"Improving Generalization in Reinforcement Learning Training Regimes for Social Robot Navigation","arxiv_id":"2308.14947","date":"2023-08-29","proceeding":null,"authors":["Adam Sigal","Hsiu-Chin Lin","AJung Moon"],"abstract":"In order for autonomous mobile robots to navigate in human spaces, they must abide by our social norms. Reinforcement learning (RL) has emerged as an effective method to train sequential decision-making policies that are able to respect these norms. However, a large portion of existing work in the field conducts both RL training and testing in simplistic environments. This limits the generalization potential of these models to unseen environments, and the meaningfulness of their reported results. We propose a method to improve the generalization performance of RL social navigation methods using curriculum learning. By employing multiple environment types and by modeling pedestrians using multiple dynamics models, we are able to progressively diversify and escalate difficulty in training. Our results show that the use of curriculum learning in training can be used to achieve better generalization performance than previous training methods. We also show that results presented in many existing state-of-the-art RL social navigation works do not evaluate their methods outside of their training environments, and thus do not reflect their policies' failure to adequately generalize to out-of-distribution scenarios. In response, we validate our training approach on larger and more crowded testing environments than those used in training, allowing for more meaningful measurements of model performance.","url_abs":"https://arxiv.org/abs/2308.14947v2","url_pdf":"https://arxiv.org/pdf/2308.14947v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"improving-reinforcement-learning-training","repo_url":"https://github.com/raise-lab/soc-nav-training","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"navigate","task_name":"Navigate"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"robot-navigation","task_name":"Robot Navigation"},{"task_slug":"sequential-decision-making","task_name":"Sequential Decision Making"},{"task_slug":"social-navigation","task_name":"Social Navigation"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2308.14947","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2308.14947"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/RAISE-Lab/soc-nav-training","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/raise-lab/soc-nav-training","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"14b63195abc494ce","entry":"average","repo":"RAISE-Lab/soc-nav-training","repo_kind":"official","path":"CrowdNav/crowd_nav/utils/explorer.py","file_url":"https://github.com/RAISE-Lab/soc-nav-training/blob/HEAD/CrowdNav/crowd_nav/utils/explorer.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"14b63195abc494ce"}},{"code_sha256_prefix":"41b3a883a96e7454","entry":"mlp","repo":"RAISE-Lab/soc-nav-training","repo_kind":"official","path":"CrowdNav/crowd_nav/policy/cadrl.py","file_url":"https://github.com/RAISE-Lab/soc-nav-training/blob/HEAD/CrowdNav/crowd_nav/policy/cadrl.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"41b3a883a96e7454"}},{"code_sha256_prefix":"36abda460d4a8430","entry":"running_mean","repo":"RAISE-Lab/soc-nav-training","repo_kind":"official","path":"CrowdNav/crowd_nav/utils/plot.py","file_url":"https://github.com/RAISE-Lab/soc-nav-training/blob/HEAD/CrowdNav/crowd_nav/utils/plot.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"36abda460d4a8430"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}