{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-video-generation-prediction-and","title":"Deep Video Generation, Prediction and Completion of Human Action Sequences","arxiv_id":"1711.08682","date":"2017-11-23","proceeding":"ECCV 2018 9","authors":["Haoye Cai","Chunyan Bai","Yu-Wing Tai","Chi-Keung Tang"],"abstract":"Current deep learning results on video generation are limited while there are\nonly a few first results on video prediction and no relevant significant\nresults on video completion. This is due to the severe ill-posedness inherent\nin these three problems. In this paper, we focus on human action videos, and\npropose a general, two-stage deep framework to generate human action videos\nwith no constraints or arbitrary number of constraints, which uniformly address\nthe three problems: video generation given no input frames, video prediction\ngiven the first few frames, and video completion given the first and last\nframes. To make the problem tractable, in the first stage we train a deep\ngenerative model that generates a human pose sequence from random noise. In the\nsecond stage, a skeleton-to-image network is trained, which is used to generate\na human action video given the complete human pose sequence generated in the\nfirst stage. By introducing the two-stage strategy, we sidestep the original\nill-posed problems while producing for the first time high-quality video\ngeneration/prediction/completion results of much longer duration. We present\nquantitative and qualitative evaluation to show that our two-stage approach\noutperforms state-of-the-art methods in video generation, prediction and video\ncompletion. Our video result demonstration can be viewed at\nhttps://iamacewhite.github.io/supp/index.html","url_abs":"http://arxiv.org/abs/1711.08682v3","url_pdf":"http://arxiv.org/pdf/1711.08682v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"human-action-generation","task_name":"Human action generation"},{"task_slug":"prediction","task_name":"Prediction"},{"task_slug":"video-generation","task_name":"Video Generation"},{"task_slug":"video-prediction","task_name":"Video Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/human-action-generation-on-human3-6m","task":"Human action generation","dataset":"Human3.6M","model":"Deep Video Generation, Prediction and Completion of Human Action Sequences","rank_in_archive_order":5,"of":5,"metrics":{"MMDa":"0.419","MMDs":"0.436"},"uses_additional_data":false},{"leaderboard":"/sota/human-action-generation-on-ntu-rgb-d-2d","task":"Human action generation","dataset":"NTU RGB+D 2D","model":"SkeletonGAN","rank_in_archive_order":5,"of":5,"metrics":{"MMDa (CS)":"0.698","MMDa (CV)":"0.999","MMDs (CS)":"0.788","MMDs (CV)":"1.311"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1711.08682","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}