{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-note-on-the-evaluation-of-generative-models","title":"A note on the evaluation of generative models","arxiv_id":"1511.01844","date":"2015-11-05","proceeding":null,"authors":["Lucas Theis","Aäron van den Oord","Matthias Bethge"],"abstract":"Probabilistic generative models can be used for compression, denoising,\ninpainting, texture synthesis, semi-supervised learning, unsupervised feature\nlearning, and other tasks. Given this wide range of applications, it is not\nsurprising that a lot of heterogeneity exists in the way these models are\nformulated, trained, and evaluated. As a consequence, direct comparison between\nmodels is often difficult. This article reviews mostly known but often\nunderappreciated properties relating to the evaluation and interpretation of\ngenerative models with a focus on image models. In particular, we show that\nthree of the currently most commonly used criteria---average log-likelihood,\nParzen window estimates, and visual fidelity of samples---are largely\nindependent of each other when the data is high-dimensional. Good performance\nwith respect to one criterion therefore need not imply good performance with\nrespect to the other criteria. Our results show that extrapolation from one\ncriterion to another is not warranted and generative models need to be\nevaluated directly with respect to the application(s) they were intended for.\nIn addition, we provide examples demonstrating that Parzen window estimates\nshould generally be avoided.","url_abs":"http://arxiv.org/abs/1511.01844v3","url_pdf":"http://arxiv.org/pdf/1511.01844v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-note-on-the-evaluation-of-generative-models","repo_url":"https://github.com/kpandey008/DCGANS","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"denoising","task_name":"Denoising"},{"task_slug":"texture-synthesis","task_name":"Texture Synthesis"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1511.01844","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}