{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/transformation-based-adversarial-video","title":"Transformation-based Adversarial Video Prediction on Large-Scale Data","arxiv_id":"2003.04035","date":"2020-03-09","proceeding":null,"authors":["Pauline Luc","Aidan Clark","Sander Dieleman","Diego de Las Casas","Yotam Doron","Albin Cassirer","Karen Simonyan"],"abstract":"Recent breakthroughs in adversarial generative modeling have led to models capable of producing video samples of high quality, even on large and complex datasets of real-world video. In this work, we focus on the task of video prediction, where given a sequence of frames extracted from a video, the goal is to generate a plausible future sequence. We first improve the state of the art by performing a systematic empirical study of discriminator decompositions and proposing an architecture that yields faster convergence and higher performance than previous approaches. We then analyze recurrent units in the generator, and propose a novel recurrent unit which transforms its past hidden state according to predicted motion-like features, and refines it to handle dis-occlusions, scene changes and other complex behavior. We show that this recurrent unit consistently outperforms previous designs. Our final model leads to a leap in the state-of-the-art performance, obtaining a test set Frechet Video Distance of 25.7, down from 69.2, on the large-scale Kinetics-600 dataset.","url_abs":"https://arxiv.org/abs/2003.04035v3","url_pdf":"https://arxiv.org/pdf/2003.04035v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"prediction","task_name":"Prediction"},{"task_slug":"video-generation","task_name":"Video Generation"},{"task_slug":"video-prediction","task_name":"Video Prediction"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"3d-convolution","method_name":"3D Convolution"},{"method_slug":"adam","method_name":"Adam"},{"method_slug":"average-pooling","method_name":"Average Pooling"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"conditional-batch-normalization","method_name":"Conditional Batch Normalization"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dvd-gan-dblock","method_name":"DVD-GAN DBlock"},{"method_slug":"dvd-gan-gblock","method_name":"DVD-GAN GBlock"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"early-stopping","method_name":"Early Stopping"},{"method_slug":"feedforward-network","method_name":"Feedforward Network"},{"method_slug":"gan-hinge-loss","method_name":"GAN Hinge Loss"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"orthogonal-regularization","method_name":"Orthogonal Regularization"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"spectral-normalization","method_name":"Spectral Normalization"},{"method_slug":"tsruc","method_name":"TSRUc"},{"method_slug":"tsrup","method_name":"TSRUp"},{"method_slug":"tsrus","method_name":"TSRUs"},{"method_slug":"ttur","method_name":"TTUR"},{"method_slug":"trivd-gan","method_name":"TrIVD-GAN"}],"datasets_introduced":[],"methods_introduced":[{"slug":"tsruc","name":"TSRUc","full_name":"TSRUc"},{"slug":"tsrup","name":"TSRUp","full_name":"TSRUp"},{"slug":"tsrus","name":"TSRUs","full_name":"TSRUs"}],"results":[{"leaderboard":"/sota/video-generation-on-bair-robot-pushing","task":"Video Generation","dataset":"BAIR Robot Pushing","model":"TrIVD-GAN-FP","rank_in_archive_order":10,"of":31,"metrics":{"Cond":"1","FVD score":"103.3","Pred":"15","Train":"15"},"uses_additional_data":false},{"leaderboard":"/sota/video-prediction-on-kinetics-600-12-frames","task":"Video Prediction","dataset":"Kinetics-600 12 frames, 64x64","model":"TriVD-GAN-FP","rank_in_archive_order":10,"of":16,"metrics":{"Cond":"5","FVD":"25.74±0.66","IS":"12.54±0.06","Pred":"11"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2003.04035","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}