{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mode-seeking-generative-adversarial-networks","title":"Mode Seeking Generative Adversarial Networks for Diverse Image Synthesis","arxiv_id":"1903.05628","date":"2019-03-13","proceeding":"CVPR 2019 6","authors":["Qi Mao","Hsin-Ying Lee","Hung-Yu Tseng","Siwei Ma","Ming-Hsuan Yang"],"abstract":"Most conditional generation tasks expect diverse outputs given a single conditional context. However, conditional generative adversarial networks (cGANs) often focus on the prior conditional information and ignore the input noise vectors, which contribute to the output variations. Recent attempts to resolve the mode collapse issue for cGANs are usually task-specific and computationally expensive. In this work, we propose a simple yet effective regularization term to address the mode collapse issue for cGANs. The proposed method explicitly maximizes the ratio of the distance between generated images with respect to the corresponding latent codes, thus encouraging the generators to explore more minor modes during training. This mode seeking regularization term is readily applicable to various conditional generation tasks without imposing training overhead or modifying the original network structures. We validate the proposed algorithm on three conditional image synthesis tasks including categorical generation, image-to-image translation, and text-to-image synthesis with different baseline models. Both qualitative and quantitative results demonstrate the effectiveness of the proposed regularization method for improving diversity without loss of quality.","url_abs":"https://arxiv.org/abs/1903.05628v6","url_pdf":"https://arxiv.org/pdf/1903.05628v6.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mode-seeking-generative-adversarial-networks","repo_url":"https://github.com/HelenMao/MSGAN","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"mode-seeking-generative-adversarial-networks","repo_url":"https://github.com/JPlin/MSGAN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"diversity","task_name":"Diversity"},{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"image-to-image-translation","task_name":"Image-to-Image Translation"},{"task_slug":"multimodal-unsupervised-image-to-image","task_name":"Multimodal Unsupervised Image-To-Image Translation"},{"task_slug":"translation","task_name":"Translation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-generation-on-cifar-10","task":"Image Generation","dataset":"CIFAR-10","model":"MSGAN","rank_in_archive_order":67,"of":78,"metrics":{"FID":"28.73"},"uses_additional_data":false},{"leaderboard":"/sota/multimodal-unsupervised-image-to-image-5","task":"Multimodal Unsupervised Image-To-Image Translation","dataset":"AFHQ","model":"MSGAN","rank_in_archive_order":3,"of":4,"metrics":{"FID":"61.4"},"uses_additional_data":false},{"leaderboard":"/sota/multimodal-unsupervised-image-to-image-4","task":"Multimodal Unsupervised Image-To-Image Translation","dataset":"CelebA-HQ","model":"MSGAN","rank_in_archive_order":3,"of":4,"metrics":{"FID":"33.1"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1903.05628","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}