{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/get-it-in-writing-formal-contracts-mitigate","title":"Formal Contracts Mitigate Social Dilemmas in Multi-Agent RL","arxiv_id":"2208.10469","date":"2022-08-22","proceeding":null,"authors":["Andreas A. Haupt","Phillip J. K. Christoffersen","Mehul Damani","Dylan Hadfield-Menell"],"abstract":"Multi-agent Reinforcement Learning (MARL) is a powerful tool for training autonomous agents acting independently in a common environment. However, it can lead to sub-optimal behavior when individual incentives and group incentives diverge. Humans are remarkably capable at solving these social dilemmas. It is an open problem in MARL to replicate such cooperative behaviors in selfish agents. In this work, we draw upon the idea of formal contracting from economics to overcome diverging incentives between agents in MARL. We propose an augmentation to a Markov game where agents voluntarily agree to binding transfers of reward, under pre-specified conditions. Our contributions are theoretical and empirical. First, we show that this augmentation makes all subgame-perfect equilibria of all Fully Observable Markov Games exhibit socially optimal behavior, given a sufficiently rich space of contracts. Next, we show that for general contract spaces, and even under partial observability, richer contract spaces lead to higher welfare. Hence, contract space design solves an exploration-exploitation tradeoff, sidestepping incentive issues. We complement our theoretical analysis with experiments. Issues of exploration in the contracting augmentation are mitigated using a training methodology inspired by multi-objective reinforcement learning: Multi-Objective Contract Augmentation Learning (MOCA). We test our methodology in static, single-move games, as well as dynamic domains that simulate traffic, pollution management and common pool resource management.","url_abs":"https://arxiv.org/abs/2208.10469v4","url_pdf":"https://arxiv.org/pdf/2208.10469v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"get-it-in-writing-formal-contracts-mitigate","repo_url":"https://github.com/algorithmic-alignment-lab/contracts","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"management","task_name":"Management"},{"task_slug":"multi-objective-reinforcement-learning","task_name":"Multi-Objective Reinforcement Learning"},{"task_slug":"multi-agent-reinforcement-learning","task_name":"Multi-agent Reinforcement Learning"},{"task_slug":"reinforcement-learning-2","task_name":"reinforcement-learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2208.10469","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2208.10469"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/algorithmic-alignment-lab/contracts","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":4},"by_repo_kind":{"official":{"samples":4,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"185bfd8a232426ff","entry":"pad_if_needed","repo":"algorithmic-alignment-lab/contracts","repo_kind":"official","path":"environments/env_utils.py","file_url":"https://github.com/algorithmic-alignment-lab/contracts/blob/HEAD/environments/env_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"185bfd8a232426ff"}},{"code_sha256_prefix":"0866f9a384a40763","entry":"pad_matrix","repo":"algorithmic-alignment-lab/contracts","repo_kind":"official","path":"environments/env_utils.py","file_url":"https://github.com/algorithmic-alignment-lab/contracts/blob/HEAD/environments/env_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0866f9a384a40763"}},{"code_sha256_prefix":"88a2d1e6de3f1713","entry":"ppo_learning","repo":"algorithmic-alignment-lab/contracts","repo_kind":"official","path":"run_training.py","file_url":"https://github.com/algorithmic-alignment-lab/contracts/blob/HEAD/run_training.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"88a2d1e6de3f1713"}},{"code_sha256_prefix":"e751c5fe02182b30","entry":"return_view","repo":"algorithmic-alignment-lab/contracts","repo_kind":"official","path":"environments/env_utils.py","file_url":"https://github.com/algorithmic-alignment-lab/contracts/blob/HEAD/environments/env_utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e751c5fe02182b30"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}