{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-distinguishability-criteria-for-estimating","title":"On distinguishability criteria for estimating generative models","arxiv_id":"1412.6515","date":"2014-12-19","proceeding":null,"authors":["Ian J. Goodfellow"],"abstract":"Two recently introduced criteria for estimation of generative models are both\nbased on a reduction to binary classification. Noise-contrastive estimation\n(NCE) is an estimation procedure in which a generative model is trained to be\nable to distinguish data samples from noise samples. Generative adversarial\nnetworks (GANs) are pairs of generator and discriminator networks, with the\ngenerator network learning to generate samples by attempting to fool the\ndiscriminator network into believing its samples are real data. Both estimation\nprocedures use the same function to drive learning, which naturally raises\nquestions about how they are related to each other, as well as whether this\nfunction is related to maximum likelihood estimation (MLE). NCE corresponds to\ntraining an internal data model belonging to the {\\em discriminator} network\nbut using a fixed generator network. We show that a variant of NCE, with a\ndynamic generator network, is equivalent to maximum likelihood estimation.\nSince pairing a learned discriminator with an appropriate dynamically selected\ngenerator recovers MLE, one might expect the reverse to hold for pairing a\nlearned generator with a certain discriminator. However, we show that\nrecovering MLE for a learned generator requires departing from the\ndistinguishability game. Specifically:\n  (i) The expected gradient of the NCE discriminator can be made to match the\nexpected gradient of\n  MLE, if one is allowed to use a non-stationary noise distribution for NCE,\n  (ii) No choice of discriminator network can make the expected gradient for\nthe GAN generator match that of MLE, and\n  (iii) The existing theory does not guarantee that GANs will converge in the\nnon-convex case.\n  This suggests that the key next step in GAN research is to determine whether\nGANs converge, and if not, to modify their training algorithm to force\nconvergence.","url_abs":"http://arxiv.org/abs/1412.6515v4","url_pdf":"http://arxiv.org/pdf/1412.6515v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-distinguishability-criteria-for-estimating","repo_url":"https://github.com/Stonesjtu/Pytorch-NCE","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"binary-classification","task_name":"Binary Classification"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1412.6515","atlas_url":"https://app.syntology.ai/?focus=1412.6515","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}