{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/warpdrive-extremely-fast-end-to-end-deep","title":"WarpDrive: Extremely Fast End-to-End Deep Multi-Agent Reinforcement Learning on a GPU","arxiv_id":"2108.13976","date":"2021-08-31","proceeding":null,"authors":["Tian Lan","Sunil Srinivasa","Huan Wang","Stephan Zheng"],"abstract":"Deep reinforcement learning (RL) is a powerful framework to train decision-making models in complex environments. However, RL can be slow as it requires repeated interaction with a simulation of the environment. In particular, there are key system engineering bottlenecks when using RL in complex environments that feature multiple agents with high-dimensional state, observation, or action spaces. We present WarpDrive, a flexible, lightweight, and easy-to-use open-source RL framework that implements end-to-end deep multi-agent RL on a single GPU (Graphics Processing Unit), built on PyCUDA and PyTorch. Using the extreme parallelization capability of GPUs, WarpDrive enables orders-of-magnitude faster RL compared to common implementations that blend CPU simulations and GPU models. Our design runs simulations and the agents in each simulation in parallel. It eliminates data copying between CPU and GPU. It also uses a single simulation data store on the GPU that is safely updated in-place. WarpDrive provides a lightweight Python interface and flexible environment wrappers that are easy to use and extend. Together, this allows the user to easily run thousands of concurrent multi-agent simulations and train on extremely large batches of experience. Through extensive experiments, we verify that WarpDrive provides high-throughput and scales almost linearly to many agents and parallel environments. For example, WarpDrive yields 2.9 million environment steps/second with 2000 environments and 1000 agents (at least 100x higher throughput compared to a CPU implementation) in a benchmark Tag simulation. As such, WarpDrive is a fast and extensible multi-agent RL platform to significantly accelerate research and development.","url_abs":"https://arxiv.org/abs/2108.13976v3","url_pdf":"https://arxiv.org/pdf/2108.13976v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"warpdrive-extremely-fast-end-to-end-deep","repo_url":"https://github.com/salesforce/warp-drive","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"BSD-3-Clause"}},{"paper_slug":"warpdrive-extremely-fast-end-to-end-deep","repo_url":"https://github.com/mila-iqia/climate-cooperation-competition","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"warpdrive-extremely-fast-end-to-end-deep","repo_url":"https://github.com/salesforce/ai-economist","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"BSD-3-Clause"}}],"tasks":[{"task_slug":null,"task_name":"CPU"},{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"deep-reinforcement-learning","task_name":"Deep Reinforcement Learning"},{"task_slug":null,"task_name":"GPU"},{"task_slug":"multi-agent-reinforcement-learning","task_name":"Multi-agent Reinforcement Learning"},{"task_slug":"reinforcement-learning-1","task_name":"Reinforcement Learning (RL)"},{"task_slug":"tag","task_name":"TAG"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2108.13976","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2108.13976"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/salesforce/warp-drive","reach":{"status":"ok","spdx":"BSD-3-Clause"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/salesforce/ai-economist","reach":{"status":"ok","spdx":"BSD-3-Clause"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/mila-iqia/climate-cooperation-competition","reach":{"status":"ok","spdx":"NOASSERTION"}}],"summary":{"unverified":7},"by_repo_kind":{"official":{"samples":2,"ran":0,"repositories":1},"listed":{"samples":5,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"d8a1bc0618312f77","entry":"cfg_dict_from_yaml","repo":"salesforce/ai-economist","repo_kind":"listed","path":"ai_economist/real_business_cycle/experiment_utils.py","file_url":"https://github.com/salesforce/ai-economist/blob/HEAD/ai_economist/real_business_cycle/experiment_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"d8a1bc0618312f77"}},{"code_sha256_prefix":"6a2b7a164a3d1480","entry":"check_if_ep_str_policy_exists","repo":"salesforce/ai-economist","repo_kind":"listed","path":"ai_economist/real_business_cycle/train_bestresponse.py","file_url":"https://github.com/salesforce/ai-economist/blob/HEAD/ai_economist/real_business_cycle/train_bestresponse.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"6a2b7a164a3d1480"}},{"code_sha256_prefix":"9c61b6cc4b60d9af","entry":"dynamic_import","repo":"salesforce/warp-drive","repo_kind":"official","path":"warp_drive/training/models/factory.py","file_url":"https://github.com/salesforce/warp-drive/blob/HEAD/warp_drive/training/models/factory.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"9c61b6cc4b60d9af"}},{"code_sha256_prefix":"443d55b2bf005c50","entry":"generate_random_actions","repo":"salesforce/warp-drive","repo_kind":"official","path":"warp_drive/env_cpu_gpu_consistency_checker.py","file_url":"https://github.com/salesforce/warp-drive/blob/HEAD/warp_drive/env_cpu_gpu_consistency_checker.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"443d55b2bf005c50"}},{"code_sha256_prefix":"580e9b27b830eb25","entry":"hash_from_dict","repo":"salesforce/ai-economist","repo_kind":"listed","path":"ai_economist/real_business_cycle/experiment_utils.py","file_url":"https://github.com/salesforce/ai-economist/blob/HEAD/ai_economist/real_business_cycle/experiment_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"580e9b27b830eb25"}},{"code_sha256_prefix":"a5f97a16853000a0","entry":"recursive_obs_dict_to_spaces_dict","repo":"salesforce/ai-economist","repo_kind":"listed","path":"ai_economist/foundation/env_wrapper.py","file_url":"https://github.com/salesforce/ai-economist/blob/HEAD/ai_economist/foundation/env_wrapper.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"a5f97a16853000a0"}},{"code_sha256_prefix":"799c662498b9093c","entry":"seed_from_base_seed","repo":"salesforce/ai-economist","repo_kind":"listed","path":"ai_economist/real_business_cycle/experiment_utils.py","file_url":"https://github.com/salesforce/ai-economist/blob/HEAD/ai_economist/real_business_cycle/experiment_utils.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"BSD-3-Clause","inline_ok":true,"mcp_get_code":{"code_sha256":"799c662498b9093c"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}