{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unsupervised-video-to-video-translation","title":"Unsupervised Video-to-Video Translation","arxiv_id":"1806.03698","date":"2018-06-10","proceeding":"ICLR 2019 5","authors":["Dina Bashkirova","Ben Usman","Kate Saenko"],"abstract":"Unsupervised image-to-image translation is a recently proposed task of\ntranslating an image to a different style or domain given only unpaired image\nexamples at training time. In this paper, we formulate a new task of\nunsupervised video-to-video translation, which poses its own unique challenges.\nTranslating video implies learning not only the appearance of objects and\nscenes but also realistic motion and transitions between consecutive frames.We\ninvestigate the performance of per-frame video-to-video translation using\nexisting image-to-image translation networks, and propose a spatio-temporal 3D\ntranslator as an alternative solution to this problem. We evaluate our 3D\nmethod on multiple synthetic datasets, such as moving colorized digits, as well\nas the realistic segmentation-to-video GTA dataset and a new CT-to-MRI\nvolumetric images translation dataset. Our results show that frame-wise\ntranslation produces realistic results on a single frame level but\nunderperforms significantly on the scale of the whole video compared to our\nthree-dimensional translation approach, which is better able to learn the\ncomplex structure of video and motion and continuity of object appearance.","url_abs":"http://arxiv.org/abs/1806.03698v1","url_pdf":"http://arxiv.org/pdf/1806.03698v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"unsupervised-video-to-video-translation","repo_url":"https://github.com/dbash/CycleGAN3D","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"image-to-image-translation","task_name":"Image-to-Image Translation"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"unsupervised-image-to-image-translation","task_name":"Unsupervised Image-To-Image Translation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1806.03698","atlas_url":"https://app.syntology.ai/?focus=1806.03698","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}