{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/predicting-deeper-into-the-future-of-semantic","title":"Predicting Deeper into the Future of Semantic Segmentation","arxiv_id":"1703.07684","date":"2017-03-22","proceeding":"ICCV 2017 10","authors":["Pauline Luc","Natalia Neverova","Camille Couprie","Jakob Verbeek","Yann Lecun"],"abstract":"The ability to predict and therefore to anticipate the future is an important\nattribute of intelligence. It is also of utmost importance in real-time\nsystems, e.g. in robotics or autonomous driving, which depend on visual scene\nunderstanding for decision making. While prediction of the raw RGB pixel values\nin future video frames has been studied in previous work, here we introduce the\nnovel task of predicting semantic segmentations of future frames. Given a\nsequence of video frames, our goal is to predict segmentation maps of not yet\nobserved video frames that lie up to a second or further in the future. We\ndevelop an autoregressive convolutional neural network that learns to\niteratively generate multiple frames. Our results on the Cityscapes dataset\nshow that directly predicting future segmentations is substantially better than\npredicting and then segmenting future RGB frames. Prediction results up to half\na second in the future are visually convincing and are much more accurate than\nthose of a baseline based on warping semantic segmentations using optical flow.","url_abs":"http://arxiv.org/abs/1703.07684v3","url_pdf":"http://arxiv.org/pdf/1703.07684v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"predicting-deeper-into-the-future-of-semantic","repo_url":"https://github.com/facebookresearch/SegmPred","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"torch","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"predicting-deeper-into-the-future-of-semantic","repo_url":"https://github.com/m-serra/action-inference-for-video-prediction-benchmarking","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"autonomous-driving","task_name":"Autonomous Driving"},{"task_slug":"decision-making","task_name":"Decision Making"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"scene-understanding","task_name":"Scene Understanding"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"video-prediction","task_name":"Video Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1703.07684","atlas_url":"https://app.syntology.ai/?focus=1703.07684","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}