{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-generative-models-with-sinkhorn","title":"Learning Generative Models with Sinkhorn Divergences","arxiv_id":"1706.00292","date":"2017-06-01","proceeding":null,"authors":["Aude Genevay","Gabriel Peyré","Marco Cuturi"],"abstract":"The ability to compare two degenerate probability distributions (i.e. two\nprobability distributions supported on two distinct low-dimensional manifolds\nliving in a much higher-dimensional space) is a crucial problem arising in the\nestimation of generative models for high-dimensional observations such as those\narising in computer vision or natural language. It is known that optimal\ntransport metrics can represent a cure for this problem, since they were\nspecifically designed as an alternative to information divergences to handle\nsuch problematic scenarios. Unfortunately, training generative machines using\nOT raises formidable computational and statistical challenges, because of (i)\nthe computational burden of evaluating OT losses, (ii) the instability and lack\nof smoothness of these losses, (iii) the difficulty to estimate robustly these\nlosses and their gradients in high dimension. This paper presents the first\ntractable computational method to train large scale generative models using an\noptimal transport loss, and tackles these three issues by relying on two key\nideas: (a) entropic smoothing, which turns the original OT loss into one that\ncan be computed using Sinkhorn fixed point iterations; (b) algorithmic\n(automatic) differentiation of these iterations. These two approximations\nresult in a robust and differentiable approximation of the OT loss with\nstreamlined GPU execution. Entropic smoothing generates a family of losses\ninterpolating between Wasserstein (OT) and Maximum Mean Discrepancy (MMD), thus\nallowing to find a sweet spot leveraging the geometry of OT and the favorable\nhigh-dimensional sample complexity of MMD which comes with unbiased gradient\nestimates. The resulting computational architecture complements nicely standard\ndeep network generative models by a stack of extra layers implementing the loss\nfunction.","url_abs":"http://arxiv.org/abs/1706.00292v3","url_pdf":"http://arxiv.org/pdf/1706.00292v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-generative-models-with-sinkhorn","repo_url":"https://github.com/Alexandre-Rio/ot_generative_models","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"learning-generative-models-with-sinkhorn","repo_url":"https://github.com/leejoonhun/feature-aligned-nbeats","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":null,"task_name":"GPU"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1706.00292","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}