{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-to-generate-long-term-future-via","title":"Learning to Generate Long-term Future via Hierarchical Prediction","arxiv_id":"1704.05831","date":"2017-04-19","proceeding":"ICML 2017 8","authors":["Ruben Villegas","Jimei Yang","Yuliang Zou","Sungryull Sohn","Xunyu Lin","Honglak Lee"],"abstract":"We propose a hierarchical approach for making long-term predictions of future\nframes. To avoid inherent compounding errors in recursive pixel-level\nprediction, we propose to first estimate high-level structure in the input\nframes, then predict how that structure evolves in the future, and finally by\nobserving a single frame from the past and the predicted high-level structure,\nwe construct the future frames without having to observe any of the pixel-level\npredictions. Long-term video prediction is difficult to perform by recurrently\nobserving the predicted frames because the small errors in pixel space\nexponentially amplify as predictions are made deeper into the future. Our\napproach prevents pixel-level error propagation from happening by removing the\nneed to observe the predicted frames. Our model is built with a combination of\nLSTM and analogy based encoder-decoder convolutional neural networks, which\nindependently predict the video structure and generate the future frames,\nrespectively. In experiments, our model is evaluated on the Human3.6M and Penn\nAction datasets on the task of long-term pixel-level video prediction of humans\nperforming actions and demonstrate significantly better results than the\nstate-of-the-art.","url_abs":"http://arxiv.org/abs/1704.05831v5","url_pdf":"http://arxiv.org/pdf/1704.05831v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-to-generate-long-term-future-via","repo_url":"https://github.com/rubenvillegas/icml2017hierchvid","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"learning-to-generate-long-term-future-via","repo_url":"https://github.com/xcyan/eccv18_mtvae","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"prediction","task_name":"Prediction"},{"task_slug":"video-prediction","task_name":"Video Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1704.05831","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1704.05831"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/xcyan/eccv18_mtvae","reach":{"status":"unanswered"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/rubenvillegas/icml2017hierchvid","reach":null}],"summary":{"ran_honours":2,"unverified":1},"by_repo_kind":{"listed":{"samples":3,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"66ea27219660a78f","entry":"inverse_transform","repo":"rubenvillegas/icml2017hierchvid","repo_kind":"listed","path":"lstm_src/train_det_rnn_h36m.py","file_url":"https://github.com/rubenvillegas/icml2017hierchvid/blob/HEAD/lstm_src/train_det_rnn_h36m.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"66ea27219660a78f"}},{"code_sha256_prefix":"574268d8d2672f71","entry":"transform","repo":"rubenvillegas/icml2017hierchvid","repo_kind":"listed","path":"lstm_src/train_det_rnn_h36m.py","file_url":"https://github.com/rubenvillegas/icml2017hierchvid/blob/HEAD/lstm_src/train_det_rnn_h36m.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"574268d8d2672f71"}},{"code_sha256_prefix":"1e39144a8a91a409","entry":"merge","repo":"rubenvillegas/icml2017hierchvid","repo_kind":"listed","path":"lstm_src/train_det_rnn_h36m.py","file_url":"https://github.com/rubenvillegas/icml2017hierchvid/blob/HEAD/lstm_src/train_det_rnn_h36m.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1e39144a8a91a409"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}