{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/end-to-end-reinforcement-learning-for-online","title":"Learning Coverage Paths in Unknown Environments with Deep Reinforcement Learning","arxiv_id":"2306.16978","date":"2023-06-29","proceeding":null,"authors":["Arvi Jonnarth","Jie Zhao","Michael Felsberg"],"abstract":"Coverage path planning (CPP) is the problem of finding a path that covers the entire free space of a confined area, with applications ranging from robotic lawn mowing to search-and-rescue. When the environment is unknown, the path needs to be planned online while mapping the environment, which cannot be addressed by offline planning methods that do not allow for a flexible path space. We investigate how suitable reinforcement learning is for this challenging problem, and analyze the involved components required to efficiently learn coverage paths, such as action space, input feature representation, neural network architecture, and reward function. We propose a computationally feasible egocentric map representation based on frontiers, and a novel reward term based on total variation to promote complete coverage. Through extensive experiments, we show that our approach surpasses the performance of both previous RL-based approaches and highly specialized methods across multiple CPP variations.","url_abs":"https://arxiv.org/abs/2306.16978v4","url_pdf":"https://arxiv.org/pdf/2306.16978v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"end-to-end-reinforcement-learning-for-online","repo_url":"https://github.com/arvijj/rl-cpp","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"BSD-3-Clause-Clear"}}],"tasks":[{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2306.16978","atlas_url":"https://app.syntology.ai/?focus=2306.16978","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2306.16978"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/arvijj/rl-cpp","reach":{"status":"ok","spdx":"BSD-3-Clause-Clear"}}],"summary":{"ran_draft_wrong":1,"unverified":1},"by_repo_kind":{"community":{"samples":2,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"9ae634c751f526d2","entry":"parse_args","repo":"efc-robot/explore-bench","repo_kind":"community","path":"exploration_benchmark/scripts/DRL/single_robot_1.py","file_url":"https://github.com/efc-robot/explore-bench/blob/HEAD/exploration_benchmark/scripts/DRL/single_robot_1.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"9ae634c751f526d2"}},{"code_sha256_prefix":"63b6017a7272e40a","entry":"get_gt","repo":"efc-robot/explore-bench","repo_kind":"community","path":"exploration_benchmark/scripts/exploration_metric_for_single_robot.py","file_url":"https://github.com/efc-robot/explore-bench/blob/HEAD/exploration_benchmark/scripts/exploration_metric_for_single_robot.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"63b6017a7272e40a"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}