{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/how-not-to-train-your-generative-model","title":"How (not) to Train your Generative Model: Scheduled Sampling, Likelihood, Adversary?","arxiv_id":"1511.05101","date":"2015-11-16","proceeding":null,"authors":["Ferenc Huszár"],"abstract":"Modern applications and progress in deep learning research have created\nrenewed interest for generative models of text and of images. However, even\ntoday it is unclear what objective functions one should use to train and\nevaluate these models. In this paper we present two contributions.\n  Firstly, we present a critique of scheduled sampling, a state-of-the-art\ntraining method that contributed to the winning entry to the MSCOCO image\ncaptioning benchmark in 2015. Here we show that despite this impressive\nempirical performance, the objective function underlying scheduled sampling is\nimproper and leads to an inconsistent learning algorithm.\n  Secondly, we revisit the problems that scheduled sampling was meant to\naddress, and present an alternative interpretation. We argue that maximum\nlikelihood is an inappropriate training objective when the end-goal is to\ngenerate natural-looking samples. We go on to derive an ideal objective\nfunction to use in this situation instead. We introduce a generalisation of\nadversarial training, and show how such method can interpolate between maximum\nlikelihood training and our ideal training objective. To our knowledge this is\nthe first theoretical analysis that explains why adversarial training tends to\nproduce samples with higher perceived quality.","url_abs":"http://arxiv.org/abs/1511.05101v1","url_pdf":"http://arxiv.org/pdf/1511.05101v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"how-not-to-train-your-generative-model","repo_url":"https://github.com/bloomberg/mixce-acl2023","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-captioning","task_name":"Image Captioning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1511.05101","atlas_url":"https://app.syntology.ai/?focus=1511.05101","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}