{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2608-26571","title":"Arrive and Survive: Scaling Safe Goal-Conditioned Policy Learning from One-Bit Failure Signals","arxiv_id":"2608.26571","date":"2026-08-27","proceeding":null,"authors":["Guopeng Li","Yiyang Duan","Yiru Jiao","Chengcheng Xu"],"abstract":"Contrastive reinforcement learning (CRL) scales effectively in goal-conditioned tasks by casting policy learning into a self-supervised contrastive objective. However, in a failure-terminated Markov decision process, established CRL considers pre-failure future goals only when constructing positive samples, without accounting for the probability mass removed by failure termination. Our theoretical analysis shows that this omission induces a systematic overestimation bias in goal-reaching values. Consequently, near-failure trajectories provide disproportionately strong supervision of success despite retaining little future occupancy. Unsafe actions can thereby be reinforced through catastrophic failure bootstrapping, leading to failed policy learning and unsustainable goal-reaching behaviours. To address this problem, we introduce two minimal yet strong corrections: mass-weighted InfoNCE corrects the overweighting of short surviving futures in critic learning, and a log-survival-mass score restores the missing survival mass in policy optimization. The resulting method, Safe Contrastive Reinforcement Learning (Safe-CRL), requires only the one-bit signal provided by failure termination to scale safe goal-conditioned policy learning. Across twelve failure-prone robot navigation and locomotion tasks, Safe-CRL consistently improves survival and substantially outperforms the Scaling-CRL baseline in goal-reaching performance. Additionally, deep Safe-CRL policies exhibit complex failure-avoidance behaviours. This study completes the CRL theory under failure termination and provides a scalable safe RL framework. The code is available via https://github.com/RomainLITUD/safe-crl.","url_abs":"https://arxiv.org/abs/2608.26571","url_pdf":"https://arxiv.org/pdf/2608.26571","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2608.26571"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/RomainLITUD/safe-crl","reach":null}],"summary":{"ran":5,"unverified":4},"by_repo_kind":{"found_in_text":{"samples":9,"ran":5,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"1dda83c567ade92a","entry":"force_square_html","repo":"RomainLITUD/safe-crl","repo_kind":"found_in_text","path":"source-code/safenav_jax/visualization_rollout.py","file_url":"https://github.com/RomainLITUD/safe-crl/blob/HEAD/source-code/safenav_jax/visualization_rollout.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1dda83c567ade92a"}},{"code_sha256_prefix":"93335e28be912aee","entry":"load_env_config","repo":"RomainLITUD/safe-crl","repo_kind":"found_in_text","path":"source-code/safenav_jax/env_config.py","file_url":"https://github.com/RomainLITUD/safe-crl/blob/HEAD/source-code/safenav_jax/env_config.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"93335e28be912aee"}},{"code_sha256_prefix":"a2733bb9d3de1642","entry":"resolve_visual_env_id","repo":"RomainLITUD/safe-crl","repo_kind":"found_in_text","path":"source-code/safenav_jax/visualization_rollout.py","file_url":"https://github.com/RomainLITUD/safe-crl/blob/HEAD/source-code/safenav_jax/visualization_rollout.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a2733bb9d3de1642"}},{"code_sha256_prefix":"22686f2f6b1ec2f4","entry":"resolved_config_dict","repo":"RomainLITUD/safe-crl","repo_kind":"found_in_text","path":"source-code/experiment_artifacts.py","file_url":"https://github.com/RomainLITUD/safe-crl/blob/HEAD/source-code/experiment_artifacts.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"22686f2f6b1ec2f4"}},{"code_sha256_prefix":"d29f06c7e9867275","entry":"unique_run_dir","repo":"RomainLITUD/safe-crl","repo_kind":"found_in_text","path":"source-code/experiment_artifacts.py","file_url":"https://github.com/RomainLITUD/safe-crl/blob/HEAD/source-code/experiment_artifacts.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d29f06c7e9867275"}},{"code_sha256_prefix":"3674f6767e66d8ab","entry":"generate_unroll","repo":"RomainLITUD/safe-crl","repo_kind":"found_in_text","path":"source-code/scalable_safe_rl/evaluator.py","file_url":"https://github.com/RomainLITUD/safe-crl/blob/HEAD/source-code/scalable_safe_rl/evaluator.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"3674f6767e66d8ab"}},{"code_sha256_prefix":"eb35eebbe52cb9b6","entry":"sample_passable_hazards","repo":"RomainLITUD/safe-crl","repo_kind":"found_in_text","path":"source-code/safenav_jax/maze_hazard_placement.py","file_url":"https://github.com/RomainLITUD/safe-crl/blob/HEAD/source-code/safenav_jax/maze_hazard_placement.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"eb35eebbe52cb9b6"}},{"code_sha256_prefix":"2f52b222ed7014be","entry":"sync_mocap_pipeline_state_for_render","repo":"RomainLITUD/safe-crl","repo_kind":"found_in_text","path":"source-code/safenav_jax/visualization_rollout.py","file_url":"https://github.com/RomainLITUD/safe-crl/blob/HEAD/source-code/safenav_jax/visualization_rollout.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"2f52b222ed7014be"}},{"code_sha256_prefix":"cd323987e1eead1e","entry":"write_resolved_config","repo":"RomainLITUD/safe-crl","repo_kind":"found_in_text","path":"source-code/experiment_artifacts.py","file_url":"https://github.com/RomainLITUD/safe-crl/blob/HEAD/source-code/experiment_artifacts.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"cd323987e1eead1e"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"cs.LG","source":"arxiv_2026.jsonl"},"syntology_extracted_results":null}