Papers › Deep Video Generation, Prediction and Completion of Human Action Sequences

Deep Video Generation, Prediction and Completion of Human Action Sequences

23 Nov 2017ECCV 2018 9arXiv:1711.08682archive 2025-07-28

Haoye Cai, Chunyan Bai, Yu-Wing Tai, Chi-Keung Tang

Current deep learning results on video generation are limited while there are only a few first results on video prediction and no relevant significant results on video completion. This is due to the severe ill-posedness inherent in these three problems. In this paper, we focus on human action videos, and propose a general, two-stage deep framework to generate human action videos with no constraints or arbitrary number of constraints, which uniformly address the three problems: video generation given no input frames, video prediction given the first few frames, and video completion given the first and last frames. To make the problem tractable, in the first stage we train a deep generative model that generates a human pose sequence from random noise. In the second stage, a skeleton-to-image network is trained, which is used to generate a human action video given the complete human pose sequence generated in the first stage. By introducing the two-stage strategy, we sidestep the original ill-posed problems while producing for the first time high-quality video generation/prediction/completion results of much longer duration. We present quantitative and qualitative evaluation to show that our two-stage approach outperforms state-of-the-art methods in video generation, prediction and video completion. Our video result demonstration can be viewed at https://iamacewhite.github.io/supp/index.html

PaperPDFConference PDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Human action generationPredictionVideo GenerationVideo Prediction

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Human action generation Human3.6M Deep Video Generation, Prediction and Completion of Human Action Sequences MMDa 0.419 #5 of 5 Archive leaderboard report
Human action generation Human3.6M Deep Video Generation, Prediction and Completion of Human Action Sequences MMDs 0.436 #5 of 5 Archive leaderboard report
Human action generation NTU RGB+D 2D SkeletonGAN MMDa (CS) 0.698 #5 of 5 Archive leaderboard report
Human action generation NTU RGB+D 2D SkeletonGAN MMDa (CV) 0.999 #5 of 5 Archive leaderboard report
Human action generation NTU RGB+D 2D SkeletonGAN MMDs (CS) 0.788 #5 of 5 Archive leaderboard report
Human action generation NTU RGB+D 2D SkeletonGAN MMDs (CV) 1.311 #5 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections