{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/simulation-to-scaled-city-zero-shot-policy","title":"Simulation to Scaled City: Zero-Shot Policy Transfer for Traffic Control via Autonomous Vehicles","arxiv_id":"1812.06120","date":"2018-12-14","proceeding":null,"authors":["Kathy Jang","Eugene Vinitsky","Behdad Chalaki","Ben Remer","Logan Beaver","Andreas Malikopoulos","Alexandre Bayen"],"abstract":"Using deep reinforcement learning, we train control policies for autonomous\nvehicles leading a platoon of vehicles onto a roundabout. Using Flow, a library\nfor deep reinforcement learning in micro-simulators, we train two policies, one\npolicy with noise injected into the state and action space and one without any\ninjected noise. In simulation, the autonomous vehicle learns an emergent\nmetering behavior for both policies in which it slows to allow for smoother\nmerging. We then directly transfer this policy without any tuning to the\nUniversity of Delaware Scaled Smart City (UDSSC), a 1:25 scale testbed for\nconnected and automated vehicles. We characterize the performance of both\npolicies on the scaled city. We show that the noise-free policy winds up\ncrashing and only occasionally metering. However, the noise-injected policy\nconsistently performs the metering behavior and remains collision-free,\nsuggesting that the noise helps with the zero-shot policy transfer.\nAdditionally, the transferred, noise-injected policy leads to a 5% reduction of\naverage travel time and a reduction of 22% in maximum travel time in the UDSSC.\nVideos of the controllers can be found at\nhttps://sites.google.com/view/iccps-policy-transfer.","url_abs":"http://arxiv.org/abs/1812.06120v2","url_pdf":"http://arxiv.org/pdf/1812.06120v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"simulation-to-scaled-city-zero-shot-policy","repo_url":"https://github.com/flow-project/flow","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"autonomous-vehicles","task_name":"Autonomous Vehicles"},{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1812.06120","atlas_url":"https://app.syntology.ai/?focus=1812.06120","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}