{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-pose-knows-video-forecasting-by","title":"The Pose Knows: Video Forecasting by Generating Pose Futures","arxiv_id":"1705.00053","date":"2017-04-28","proceeding":"ICCV 2017 10","authors":["Jacob Walker","Kenneth Marino","Abhinav Gupta","Martial Hebert"],"abstract":"Current approaches in video forecasting attempt to generate videos directly\nin pixel space using Generative Adversarial Networks (GANs) or Variational\nAutoencoders (VAEs). However, since these approaches try to model all the\nstructure and scene dynamics at once, in unconstrained settings they often\ngenerate uninterpretable results. Our insight is to model the forecasting\nproblem at a higher level of abstraction. Specifically, we exploit human pose\ndetectors as a free source of supervision and break the video forecasting\nproblem into two discrete steps. First we explicitly model the high level\nstructure of active objects in the scene---humans---and use a VAE to model the\npossible future movements of humans in the pose space. We then use the future\nposes generated as conditional information to a GAN to predict the future\nframes of the video in pixel space. By using the structured space of pose as an\nintermediate representation, we sidestep the problems that GANs have in\ngenerating video pixels directly. We show through quantitative and qualitative\nevaluation that our method outperforms state-of-the-art methods for video\nprediction.","url_abs":"http://arxiv.org/abs/1705.00053v1","url_pdf":"http://arxiv.org/pdf/1705.00053v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-pose-knows-video-forecasting-by","repo_url":"https://github.com/KMarino/MMD_evalcode","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"human-pose-forecasting","task_name":"Human Pose Forecasting"},{"task_slug":"video-forecasting","task_name":"Video Forecasting"},{"task_slug":"video-prediction","task_name":"Video Prediction"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/human-pose-forecasting-on-amass","task":"Human Pose Forecasting","dataset":"AMASS","model":"ThePoseKnows","rank_in_archive_order":11,"of":11,"metrics":{"ADE":"0.656","APD":"9.283","FDE":"0.675"},"uses_additional_data":false},{"leaderboard":"/sota/human-pose-forecasting-on-human36m","task":"Human Pose Forecasting","dataset":"Human3.6M","model":"Pose-Knows","rank_in_archive_order":30,"of":33,"metrics":{"ADE":"461","APD":"6723","CMD":"6.326","FDE":"560","FID":"0.538","MMADE":"522","MMFDE":"569"},"uses_additional_data":false},{"leaderboard":"/sota/human-pose-forecasting-on-humaneva-i","task":"Human Pose Forecasting","dataset":"HumanEva-I","model":"Pose-Knows","rank_in_archive_order":8,"of":11,"metrics":{"ADE@2000ms":"269","APD@2000ms":"2308","FDE@2000ms":"296"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1705.00053","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}