{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/adversarial-training-as-stackelberg-game-an","title":"Adversarial Regularization as Stackelberg Game: An Unrolled Optimization Approach","arxiv_id":"2104.04886","date":"2021-04-11","proceeding":"EMNLP 2021 11","authors":["Simiao Zuo","Chen Liang","Haoming Jiang","Xiaodong Liu","Pengcheng He","Jianfeng Gao","Weizhu Chen","Tuo Zhao"],"abstract":"Adversarial regularization has been shown to improve the generalization performance of deep learning models in various natural language processing tasks. Existing works usually formulate the method as a zero-sum game, which is solved by alternating gradient descent/ascent algorithms. Such a formulation treats the adversarial and the defending players equally, which is undesirable because only the defending player contributes to the generalization performance. To address this issue, we propose Stackelberg Adversarial Regularization (SALT), which formulates adversarial regularization as a Stackelberg game. This formulation induces a competition between a leader and a follower, where the follower generates perturbations, and the leader trains the model subject to the perturbations. Different from conventional approaches, in SALT, the leader is in an advantageous position. When the leader moves, it recognizes the strategy of the follower and takes the anticipated follower's outcomes into consideration. Such a leader's advantage enables us to improve the model fitting to the unperturbed data. The leader's strategic information is captured by the Stackelberg gradient, which is obtained using an unrolling algorithm. Our experimental results on a set of machine translation and natural language understanding tasks show that SALT outperforms existing adversarial regularization baselines across all tasks. Our code is available at https://github.com/SimiaoZuo/Stackelberg-Adv.","url_abs":"https://arxiv.org/abs/2104.04886v3","url_pdf":"https://arxiv.org/pdf/2104.04886v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"adversarial-training-as-stackelberg-game-an","repo_url":"https://github.com/SimiaoZuo/Stackelberg-Adv","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"machine-translation","task_name":"Machine Translation"},{"task_slug":"natural-language-understanding","task_name":"Natural Language Understanding"},{"task_slug":"unrolling","task_name":"Rolling Shutter Correction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2104.04886","atlas_url":"https://app.syntology.ai/?focus=2104.04886","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2104.04886"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/microsoft/MT-DNN","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":2,"ran_honours":1},"by_repo_kind":{"found_in_text":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"342da45925662575","entry":"bertgelu","repo":"microsoft/MT-DNN","repo_kind":"found_in_text","path":"mtdnn/common/activation_functions.py","file_url":"https://github.com/microsoft/MT-DNN/blob/HEAD/mtdnn/common/activation_functions.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"342da45925662575"}},{"code_sha256_prefix":"691b566282a42d88","entry":"linear","repo":"microsoft/MT-DNN","repo_kind":"found_in_text","path":"mtdnn/common/activation_functions.py","file_url":"https://github.com/microsoft/MT-DNN/blob/HEAD/mtdnn/common/activation_functions.py","link_basis":"harvester_set","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"691b566282a42d88"}},{"code_sha256_prefix":"d1453c3ac3a28807","entry":"swish","repo":"microsoft/MT-DNN","repo_kind":"found_in_text","path":"mtdnn/common/activation_functions.py","file_url":"https://github.com/microsoft/MT-DNN/blob/HEAD/mtdnn/common/activation_functions.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d1453c3ac3a28807"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}