Papers › Exploring Spatial-Temporal Multi-Frequency Analysis for High-Fidelity and...

Exploring Spatial-Temporal Multi-Frequency Analysis for High-Fidelity and Temporal-Consistency Video Prediction

23 Feb 2020CVPR 2020 6arXiv:2002.09905archive 2025-07-28

Beibei Jin, Yu Hu, Qiankun Tang, Jingyu Niu, Zhiping Shi, Yinhe Han, Xiaowei Li

Video prediction is a pixel-wise dense prediction task to infer future frames based on past frames. Missing appearance details and motion blur are still two major problems for current predictive models, which lead to image distortion and temporal inconsistency. In this paper, we point out the necessity of exploring multi-frequency analysis to deal with the two problems. Inspired by the frequency band decomposition characteristic of Human Vision System (HVS), we propose a video prediction network based on multi-level wavelet analysis to deal with spatial and temporal information in a unified manner. Specifically, the multi-level spatial discrete wavelet transform decomposes each video frame into anisotropic sub-bands with multiple frequencies, helping to enrich structural information and reserve fine details. On the other hand, multi-level temporal discrete wavelet transform which operates on time axis decomposes the frame sequence into sub-band groups of different frequencies to accurately capture multi-frequency motions under a fixed frame rate. Extensive experiments on diverse datasets demonstrate that our model shows significant improvements on fidelity and temporal consistency over state-of-the-art works.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

Bei-Jin/STMFANet officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

PredictionVideo GenerationVideo Prediction

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Generation BAIR Robot Pushing WAM Cond 2 #20 of 31 Archive leaderboard report
Video Generation BAIR Robot Pushing WAM FVD score 159.6 #20 of 31 Archive leaderboard report
Video Generation BAIR Robot Pushing WAM LPIPS 0.0936 #20 of 31 Archive leaderboard report
Video Generation BAIR Robot Pushing WAM PSNR 21.02 #20 of 31 Archive leaderboard report
Video Generation BAIR Robot Pushing WAM Pred 28 #20 of 31 Archive leaderboard report
Video Generation BAIR Robot Pushing WAM SSIM 0.844 #20 of 31 Archive leaderboard report
Video Generation BAIR Robot Pushing WAM Train 14 #20 of 31 Archive leaderboard report
Video Prediction KTH WAM Cond 10 #14 of 31 Archive leaderboard report
Video Prediction KTH WAM PSNR 29.85 #14 of 31 Archive leaderboard report
Video Prediction KTH WAM Pred 20 #14 of 31 Archive leaderboard report
Video Prediction KTH WAM SSIM 0.893 #14 of 31 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections