Papers › The Pose Knows: Video Forecasting by Generating Pose Futures

The Pose Knows: Video Forecasting by Generating Pose Futures

28 Apr 2017ICCV 2017 10arXiv:1705.00053archive 2025-07-28

Jacob Walker, Kenneth Marino, Abhinav Gupta, Martial Hebert

Current approaches in video forecasting attempt to generate videos directly in pixel space using Generative Adversarial Networks (GANs) or Variational Autoencoders (VAEs). However, since these approaches try to model all the structure and scene dynamics at once, in unconstrained settings they often generate uninterpretable results. Our insight is to model the forecasting problem at a higher level of abstraction. Specifically, we exploit human pose detectors as a free source of supervision and break the video forecasting problem into two discrete steps. First we explicitly model the high level structure of active objects in the scene---humans---and use a VAE to model the possible future movements of humans in the pose space. We then use the future poses generated as conditional information to a GAN to predict the future frames of the video in pixel space. By using the structured space of pose as an intermediate representation, we sidestep the problems that GANs have in generating video pixels directly. We show through quantitative and qualitative evaluation that our method outperforms state-of-the-art methods for video prediction.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Human Pose ForecastingVideo ForecastingVideo Prediction

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Human Pose Forecasting AMASS ThePoseKnows ADE 0.656 #11 of 11 Archive leaderboard report
Human Pose Forecasting AMASS ThePoseKnows APD 9.283 #11 of 11 Archive leaderboard report
Human Pose Forecasting AMASS ThePoseKnows FDE 0.675 #11 of 11 Archive leaderboard report
Human Pose Forecasting Human3.6M Pose-Knows ADE 461 #30 of 33 Archive leaderboard report
Human Pose Forecasting Human3.6M Pose-Knows APD 6723 #30 of 33 Archive leaderboard report
Human Pose Forecasting Human3.6M Pose-Knows CMD 6.326 #30 of 33 Archive leaderboard report
Human Pose Forecasting Human3.6M Pose-Knows FDE 560 #30 of 33 Archive leaderboard report
Human Pose Forecasting Human3.6M Pose-Knows FID 0.538 #30 of 33 Archive leaderboard report
Human Pose Forecasting Human3.6M Pose-Knows MMADE 522 #30 of 33 Archive leaderboard report
Human Pose Forecasting Human3.6M Pose-Knows MMFDE 569 #30 of 33 Archive leaderboard report
Human Pose Forecasting HumanEva-I Pose-Knows ADE@2000ms 269 #8 of 11 Archive leaderboard report
Human Pose Forecasting HumanEva-I Pose-Knows APD@2000ms 2308 #8 of 11 Archive leaderboard report
Human Pose Forecasting HumanEva-I Pose-Knows FDE@2000ms 296 #8 of 11 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Convolution

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections