{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/video-to-video-synthesis","title":"Video-to-Video Synthesis","arxiv_id":"1808.06601","date":"2018-08-20","proceeding":"NeurIPS 2018 12","authors":["Ting-Chun Wang","Ming-Yu Liu","Jun-Yan Zhu","Guilin Liu","Andrew Tao","Jan Kautz","Bryan Catanzaro"],"abstract":"We study the problem of video-to-video synthesis, whose goal is to learn a\nmapping function from an input source video (e.g., a sequence of semantic\nsegmentation masks) to an output photorealistic video that precisely depicts\nthe content of the source video. While its image counterpart, the\nimage-to-image synthesis problem, is a popular topic, the video-to-video\nsynthesis problem is less explored in the literature. Without understanding\ntemporal dynamics, directly applying existing image synthesis approaches to an\ninput video often results in temporally incoherent videos of low visual\nquality. In this paper, we propose a novel video-to-video synthesis approach\nunder the generative adversarial learning framework. Through carefully-designed\ngenerator and discriminator architectures, coupled with a spatio-temporal\nadversarial objective, we achieve high-resolution, photorealistic, temporally\ncoherent video results on a diverse set of input formats including segmentation\nmasks, sketches, and poses. Experiments on multiple benchmarks show the\nadvantage of our method compared to strong baselines. In particular, our model\nis capable of synthesizing 2K resolution videos of street scenes up to 30\nseconds long, which significantly advances the state-of-the-art of video\nsynthesis. Finally, we apply our approach to future video prediction,\noutperforming several state-of-the-art competing systems.","url_abs":"http://arxiv.org/abs/1808.06601v2","url_pdf":"http://arxiv.org/pdf/1808.06601v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"video-to-video-synthesis","repo_url":"https://github.com/NVIDIA/vid2vid","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"video-to-video-synthesis","repo_url":"https://github.com/BUTIYO/vid2vid-test","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"video-to-video-synthesis","repo_url":"https://github.com/Sjunna9819/My-First-Project","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"video-to-video-synthesis","repo_url":"https://github.com/divyanshpuri02/Nvidia","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"video-to-video-synthesis","repo_url":"https://github.com/divyanshpuri02/divyansh.github.io","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"video-to-video-synthesis","repo_url":"https://github.com/eric-erki/vid2vid","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"video-to-video-synthesis","repo_url":"https://github.com/fniroui/depth2room","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"video-to-video-synthesis","repo_url":"https://github.com/freedombenLiu/vid2vid","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"video-to-video-synthesis","repo_url":"https://github.com/sakshamgupta006/video-to-video-synthesis","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"video-to-video-synthesis","repo_url":"https://github.com/yawayo/vid2vid","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"2k","task_name":"2k"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"video-prediction","task_name":"Video Prediction"},{"task_slug":"video-deraining","task_name":"Video deraining"},{"task_slug":"video-to-video-synthesis","task_name":"Video-to-Video Synthesis"}],"methods":[],"datasets_introduced":[{"slug":"kayifamilytv-video-solution","name":"KayifamilyTv Video Solution","full_name":"KayifamilyTv Video Solution: Scalable Video-Mining Pipeline"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-deraining-on-video-waterdrop-removal","task":"Video deraining","dataset":"Video Waterdrop Removal Dataset","model":"Vid2Vid","rank_in_archive_order":4,"of":5,"metrics":{"PSNR":"28.73","SSIM":"0.9542"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1808.06601","atlas_url":"https://app.syntology.ai/?focus=1808.06601","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}