{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/what-game-are-we-playing-end-to-end-learning","title":"What game are we playing? End-to-end learning in normal and extensive form games","arxiv_id":"1805.02777","date":"2018-05-07","proceeding":null,"authors":["Chun Kai Ling","Fei Fang","J. Zico Kolter"],"abstract":"Although recent work in AI has made great progress in solving large,\nzero-sum, extensive-form games, the underlying assumption in most past work is\nthat the parameters of the game itself are known to the agents. This paper\ndeals with the relatively under-explored but equally important \"inverse\"\nsetting, where the parameters of the underlying game are not known to all\nagents, but must be learned through observations. We propose a differentiable,\nend-to-end learning framework for addressing this task. In particular, we\nconsider a regularized version of the game, equivalent to a particular form of\nquantal response equilibrium, and develop 1) a primal-dual Newton method for\nfinding such equilibrium points in both normal and extensive form games; and 2)\na backpropagation method that lets us analytically compute gradients of all\nrelevant game parameters through the solution itself. This ultimately lets us\nlearn the game by training in an end-to-end fashion, effectively by integrating\na \"differentiable game solver\" into the loop of larger deep network\narchitectures. We demonstrate the effectiveness of the learning method in\nseveral settings including poker and security game tasks.","url_abs":"http://arxiv.org/abs/1805.02777v2","url_pdf":"http://arxiv.org/pdf/1805.02777v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"what-game-are-we-playing-end-to-end-learning","repo_url":"https://github.com/lingchunkai/payoff_learning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"form","task_name":"Form"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1805.02777","atlas_url":"https://app.syntology.ai/?focus=1805.02777","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1805.02777"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lingchunkai/payoff_learning","reach":null}],"summary":{"ran_honours":1},"by_repo_kind":{"listed":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"d16cfc59de2f0951","entry":"Evaluate","repo":"lingchunkai/payoff_learning","repo_kind":"listed","path":"src/experiments/experiments.py","file_url":"https://github.com/lingchunkai/payoff_learning/blob/HEAD/src/experiments/experiments.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"d16cfc59de2f0951"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}