{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/variational-discriminator-bottleneck","title":"Variational Discriminator Bottleneck: Improving Imitation Learning, Inverse RL, and GANs by Constraining Information Flow","arxiv_id":"1810.00821","date":"2018-10-01","proceeding":"ICLR 2019 5","authors":["Xue Bin Peng","Angjoo Kanazawa","Sam Toyer","Pieter Abbeel","Sergey Levine"],"abstract":"Adversarial learning methods have been proposed for a wide range of applications, but the training of adversarial models can be notoriously unstable. Effectively balancing the performance of the generator and discriminator is critical, since a discriminator that achieves very high accuracy will produce relatively uninformative gradients. In this work, we propose a simple and general technique to constrain information flow in the discriminator by means of an information bottleneck. By enforcing a constraint on the mutual information between the observations and the discriminator's internal representation, we can effectively modulate the discriminator's accuracy and maintain useful and informative gradients. We demonstrate that our proposed variational discriminator bottleneck (VDB) leads to significant improvements across three distinct application areas for adversarial learning algorithms. Our primary evaluation studies the applicability of the VDB to imitation learning of dynamic continuous control skills, such as running. We show that our method can learn such skills directly from \\emph{raw} video demonstrations, substantially outperforming prior adversarial imitation learning methods. The VDB can also be combined with adversarial inverse reinforcement learning to learn parsimonious reward functions that can be transferred and re-optimized in new settings. Finally, we demonstrate that VDB can train GANs more effectively for image generation, improving upon a number of prior stabilization methods.","url_abs":"https://arxiv.org/abs/1810.00821v4","url_pdf":"https://arxiv.org/pdf/1810.00821v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"variational-discriminator-bottleneck","repo_url":"https://github.com/akanimax/Variational_Discriminator_Bottleneck","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"variational-discriminator-bottleneck","repo_url":"https://github.com/createamind/VDB-GAN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"variational-discriminator-bottleneck","repo_url":"https://github.com/naruya/vgan-pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"variational-discriminator-bottleneck","repo_url":"https://github.com/reinforcement-learning-kr/lets-do-irl","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"variational-discriminator-bottleneck","repo_url":"https://github.com/shamanez/Variational-Discriminator-Bottleneck-Tensorflow-Implementation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"continuous-control","task_name":"Continuous Control"},{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"imitation-learning","task_name":"Imitation Learning"},{"task_slug":"reinforcement-learning","task_name":"Reinforcement Learning"},{"task_slug":"continuous-control","task_name":"continuous-control"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1810.00821","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1810.00821"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/shamanez/Variational-Discriminator-Bottleneck-Tensorflow-Implementation","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/naruya/vgan-pytorch","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/akanimax/Variational_Discriminator_Bottleneck","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/createamind/VDB-GAN","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/reinforcement-learning-kr/lets-do-irl","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":13},"by_repo_kind":{"listed":{"samples":13,"ran":0,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"b065405787f332a5","entry":"actvn","repo":"akanimax/Variational_Discriminator_Bottleneck","repo_kind":"listed","path":"source/vdb/Gan_networks.py","file_url":"https://github.com/akanimax/Variational_Discriminator_Bottleneck/blob/HEAD/source/vdb/Gan_networks.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b065405787f332a5"}},{"code_sha256_prefix":"9778db32c03260b1","entry":"expert_feature_expectations","repo":"reinforcement-learning-kr/lets-do-irl","repo_kind":"listed","path":"mountaincar/maxent/maxent.py","file_url":"https://github.com/reinforcement-learning-kr/lets-do-irl/blob/HEAD/mountaincar/maxent/maxent.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9778db32c03260b1"}},{"code_sha256_prefix":"b5a95d29758c57e3","entry":"get_action","repo":"reinforcement-learning-kr/lets-do-irl","repo_kind":"listed","path":"mujoco/gail/utils/utils.py","file_url":"https://github.com/reinforcement-learning-kr/lets-do-irl/blob/HEAD/mujoco/gail/utils/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b5a95d29758c57e3"}},{"code_sha256_prefix":"4c285e4ad37fd444","entry":"get_data_loader","repo":"akanimax/Variational_Discriminator_Bottleneck","repo_kind":"listed","path":"source/data_processing/DataLoader.py","file_url":"https://github.com/akanimax/Variational_Discriminator_Bottleneck/blob/HEAD/source/data_processing/DataLoader.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4c285e4ad37fd444"}},{"code_sha256_prefix":"f80db665bb06675c","entry":"get_entropy","repo":"reinforcement-learning-kr/lets-do-irl","repo_kind":"listed","path":"mujoco/gail/utils/utils.py","file_url":"https://github.com/reinforcement-learning-kr/lets-do-irl/blob/HEAD/mujoco/gail/utils/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f80db665bb06675c"}},{"code_sha256_prefix":"e83aab097d98bcde","entry":"get_gae","repo":"reinforcement-learning-kr/lets-do-irl","repo_kind":"listed","path":"mujoco/gail/train_model.py","file_url":"https://github.com/reinforcement-learning-kr/lets-do-irl/blob/HEAD/mujoco/gail/train_model.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e83aab097d98bcde"}},{"code_sha256_prefix":"ac908e366f291aaa","entry":"get_reward","repo":"reinforcement-learning-kr/lets-do-irl","repo_kind":"listed","path":"mountaincar/maxent/maxent.py","file_url":"https://github.com/reinforcement-learning-kr/lets-do-irl/blob/HEAD/mountaincar/maxent/maxent.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"ac908e366f291aaa"}},{"code_sha256_prefix":"e36470391de100a1","entry":"get_transform","repo":"akanimax/Variational_Discriminator_Bottleneck","repo_kind":"listed","path":"source/data_processing/DataLoader.py","file_url":"https://github.com/akanimax/Variational_Discriminator_Bottleneck/blob/HEAD/source/data_processing/DataLoader.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e36470391de100a1"}},{"code_sha256_prefix":"a2dbc3ebed006ccb","entry":"log_prob_density","repo":"reinforcement-learning-kr/lets-do-irl","repo_kind":"listed","path":"mujoco/gail/utils/utils.py","file_url":"https://github.com/reinforcement-learning-kr/lets-do-irl/blob/HEAD/mujoco/gail/utils/utils.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"a2dbc3ebed006ccb"}},{"code_sha256_prefix":"d5151aa5668d7556","entry":"read_loss_log","repo":"akanimax/Variational_Discriminator_Bottleneck","repo_kind":"listed","path":"source/generate_loss_plots.py","file_url":"https://github.com/akanimax/Variational_Discriminator_Bottleneck/blob/HEAD/source/generate_loss_plots.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d5151aa5668d7556"}},{"code_sha256_prefix":"372226043ee94c22","entry":"surrogate_loss","repo":"reinforcement-learning-kr/lets-do-irl","repo_kind":"listed","path":"mujoco/vail/train_model.py","file_url":"https://github.com/reinforcement-learning-kr/lets-do-irl/blob/HEAD/mujoco/vail/train_model.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"372226043ee94c22"}},{"code_sha256_prefix":"9d37eb48f1a728b9","entry":"train_discrim","repo":"reinforcement-learning-kr/lets-do-irl","repo_kind":"listed","path":"mujoco/gail/train_model.py","file_url":"https://github.com/reinforcement-learning-kr/lets-do-irl/blob/HEAD/mujoco/gail/train_model.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"9d37eb48f1a728b9"}},{"code_sha256_prefix":"e048446774c5bd60","entry":"train_vdb","repo":"reinforcement-learning-kr/lets-do-irl","repo_kind":"listed","path":"mujoco/vail/train_model.py","file_url":"https://github.com/reinforcement-learning-kr/lets-do-irl/blob/HEAD/mujoco/vail/train_model.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e048446774c5bd60"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}