Papers › Video-to-Video Synthesis

Video-to-Video Synthesis

20 Aug 2018NeurIPS 2018 12arXiv:1808.06601archive 2025-07-28

Ting-Chun Wang, Ming-Yu Liu, Jun-Yan Zhu, Guilin Liu, Andrew Tao, Jan Kautz, Bryan Catanzaro

We study the problem of video-to-video synthesis, whose goal is to learn a mapping function from an input source video (e.g., a sequence of semantic segmentation masks) to an output photorealistic video that precisely depicts the content of the source video. While its image counterpart, the image-to-image synthesis problem, is a popular topic, the video-to-video synthesis problem is less explored in the literature. Without understanding temporal dynamics, directly applying existing image synthesis approaches to an input video often results in temporally incoherent videos of low visual quality. In this paper, we propose a novel video-to-video synthesis approach under the generative adversarial learning framework. Through carefully-designed generator and discriminator architectures, coupled with a spatio-temporal adversarial objective, we achieve high-resolution, photorealistic, temporally coherent video results on a diverse set of input formats including segmentation masks, sketches, and poses. Experiments on multiple benchmarks show the advantage of our method compared to strong baselines. In particular, our model is capable of synthesizing 2K resolution videos of street scenes up to 30 seconds long, which significantly advances the state-of-the-art of video synthesis. Finally, we apply our approach to future video prediction, outperforming several state-of-the-art competing systems.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

NVIDIA/vid2vid officialmentioned in papermentioned on GitHubpytorch report
BUTIYO/vid2vid-test mentioned on GitHubpytorch report
Sjunna9819/My-First-Project mentioned on GitHubpytorch report
divyanshpuri02/Nvidia mentioned on GitHubpytorch report
divyanshpuri02/divyansh.github.io mentioned on GitHubpytorch report
eric-erki/vid2vid mentioned on GitHubpytorch report
fniroui/depth2room mentioned on GitHubpytorch report
freedombenLiu/vid2vid mentioned on GitHubpytorch report
yawayo/vid2vid mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

2kSemantic SegmentationVideo PredictionVideo derainingVideo-to-Video Synthesis

Datasets

Introduced by this paper, per the archive.

KayifamilyTv Video Solution

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video deraining Video Waterdrop Removal Dataset Vid2Vid PSNR 28.73 #4 of 5 Archive leaderboard report
Video deraining Video Waterdrop Removal Dataset Vid2Vid SSIM 0.9542 #4 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections