Papers › MSPred: Video Prediction at Multiple Spatio-Temporal Scales with Hierarchical...

MSPred: Video Prediction at Multiple Spatio-Temporal Scales with Hierarchical Recurrent Networks

17 Mar 2022arXiv:2203.09303archive 2025-07-28

Angel Villar-Corrales, Ani Karapetyan, Andreas Boltres, Sven Behnke

Autonomous systems not only need to understand their current environment, but should also be able to predict future actions conditioned on past states, for instance based on captured camera frames. However, existing models mainly focus on forecasting future video frames for short time-horizons, hence being of limited use for long-term action planning. We propose Multi-Scale Hierarchical Prediction (MSPred), a novel video prediction model able to simultaneously forecast future possible outcomes of different levels of granularity at different spatio-temporal scales. By combining spatial and temporal downsampling, MSPred efficiently predicts abstract representations such as human poses or locations over long time horizons, while still maintaining a competitive performance for video frame prediction. In our experiments, we demonstrate that MSPred accurately predicts future video frames as well as high-level representations (e.g. keypoints or semantics) on bin-picking and action recognition datasets, while consistently outperforming popular approaches for future frame prediction. Furthermore, we ablate different modules and design choices in MSPred, experimentally validating that combining features of different spatial and temporal granularity leads to a superior performance. Code and models to reproduce our experiments can be found in https://github.com/AIS-Bonn/MSPred.

PaperPDFCode

Code

AIS-Bonn/MSPred officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

PredictionVideo Prediction

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Prediction KTH MSPred LPIPS 0.029 #13 of 31 Archive leaderboard report
Video Prediction KTH MSPred MSE 23.18 #13 of 31 Archive leaderboard report
Video Prediction KTH MSPred PSNR 27.81 #13 of 31 Archive leaderboard report
Video Prediction KTH MSPred SSIM 0.951 #13 of 31 Archive leaderboard report
Video Prediction Moving MNIST MSPred LPIPS 0.024 #21 of 31 Archive leaderboard report
Video Prediction Moving MNIST MSPred MSE 34.44 #21 of 31 Archive leaderboard report
Video Prediction Moving MNIST MSPred PSNR 26.82 #21 of 31 Archive leaderboard report
Video Prediction Moving MNIST MSPred SSIM 0.975 #21 of 31 Archive leaderboard report
Video Prediction SynpickVP MSPred LPIPS 0.033 #1 of 5 Archive leaderboard report
Video Prediction SynpickVP MSPred MSE 53.09 #1 of 5 Archive leaderboard report
Video Prediction SynpickVP MSPred PSNR 27.89 #1 of 5 Archive leaderboard report
Video Prediction SynpickVP MSPred SSIM 0.881 #1 of 5 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

ConvLSTMConvolutionSigmoid ActivationTanh Activation

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections