{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-to-generate-images-with-perceptual","title":"Learning to Generate Images with Perceptual Similarity Metrics","arxiv_id":"1511.06409","date":"2015-11-19","proceeding":null,"authors":["Jake Snell","Karl Ridgeway","Renjie Liao","Brett D. Roads","Michael C. Mozer","Richard S. Zemel"],"abstract":"Deep networks are increasingly being applied to problems involving image\nsynthesis, e.g., generating images from textual descriptions and reconstructing\nan input image from a compact representation. Supervised training of\nimage-synthesis networks typically uses a pixel-wise loss (PL) to indicate the\nmismatch between a generated image and its corresponding target image. We\npropose instead to use a loss function that is better calibrated to human\nperceptual judgments of image quality: the multiscale structural-similarity\nscore (MS-SSIM). Because MS-SSIM is differentiable, it is easily incorporated\ninto gradient-descent learning. We compare the consequences of using MS-SSIM\nversus PL loss on training deterministic and stochastic autoencoders. For three\ndifferent architectures, we collected human judgments of the quality of image\nreconstructions. Observers reliably prefer images synthesized by\nMS-SSIM-optimized models over those synthesized by PL-optimized models, for two\ndistinct PL measures ($\\ell_1$ and $\\ell_2$ distances). We also explore the\neffect of training objective on image encoding and analyze conditions under\nwhich perceptually-optimized representations yield better performance on image\nclassification. Finally, we demonstrate the superiority of\nperceptually-optimized networks for super-resolution imaging. Just as computer\nvision has advanced through the use of convolutional architectures that mimic\nthe structure of the mammalian visual system, we argue that significant\nadditional advances can be made in modeling images through the use of training\nobjectives that are well aligned to characteristics of human perception.","url_abs":"http://arxiv.org/abs/1511.06409v3","url_pdf":"http://arxiv.org/pdf/1511.06409v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-to-generate-images-with-perceptual","repo_url":"https://github.com/clementchadebec/benchmark_VAE","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"ms-ssim","task_name":"MS-SSIM"},{"task_slug":"ssim","task_name":"SSIM"},{"task_slug":"super-resolution","task_name":"Super-Resolution"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1511.06409","atlas_url":"https://app.syntology.ai/?focus=1511.06409","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}