Papers › Generating Videos with Scene Dynamics

Generating Videos with Scene Dynamics

8 Sep 2016NeurIPS 2016 12arXiv:1609.02612archive 2025-07-28

Carl Vondrick, Hamed Pirsiavash, Antonio Torralba

We capitalize on large amounts of unlabeled video in order to learn a model of scene dynamics for both video recognition tasks (e.g. action classification) and video generation tasks (e.g. future prediction). We propose a generative adversarial network for video with a spatio-temporal convolutional architecture that untangles the scene's foreground from the background. Experiments suggest this model can generate tiny videos up to a second at full frame rate better than simple baselines, and we show its utility at predicting plausible futures of static images. Moreover, experiments and visualizations show the model internally learns useful features for recognizing actions with minimal supervision, suggesting scene dynamics are a promising signal for representation learning. We believe generative video models can impact many applications in video understanding and simulation.

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Action ClassificationFuture predictionGeneral ClassificationRepresentation LearningSelf-Supervised Action RecognitionVideo GenerationVideo RecognitionVideo Understanding

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Self-Supervised Action Recognition UCF101 VideoGan (C3D) 3-fold Accuracy 52.1 #51 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 VideoGan (C3D) Frozen false #51 of 53 Archive leaderboard report
Self-Supervised Action Recognition UCF101 VideoGan (C3D) Pre-Training Dataset UCF101 #51 of 53 Archive leaderboard report
Video Generation UCF-101 16 frames, 64x64, Unconditional VGAN Inception Score 8.18 #7 of 7 Archive leaderboard report
Video Generation UCF-101 16 frames, Unconditional, Single GPU VGAN Inception Score 8.18 #7 of 7 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections