{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/predicting-future-instance-segmentation-by","title":"Predicting Future Instance Segmentation by Forecasting Convolutional Features","arxiv_id":"1803.11496","date":"2018-03-30","proceeding":"ECCV 2018 9","authors":["Pauline Luc","Camille Couprie","Yann Lecun","Jakob Verbeek"],"abstract":"Anticipating future events is an important prerequisite towards intelligent\nbehavior. Video forecasting has been studied as a proxy task towards this goal.\nRecent work has shown that to predict semantic segmentation of future frames,\nforecasting at the semantic level is more effective than forecasting RGB frames\nand then segmenting these. In this paper we consider the more challenging\nproblem of future instance segmentation, which additionally segments out\nindividual objects. To deal with a varying number of output labels per image,\nwe develop a predictive model in the space of fixed-sized convolutional\nfeatures of the Mask R-CNN instance segmentation model. We apply the \"detection\nhead'\" of Mask R-CNN on the predicted features to produce the instance\nsegmentation of future frames. Experiments show that this approach\nsignificantly improves over strong baselines based on optical flow and\nrepurposed instance segmentation architectures.","url_abs":"http://arxiv.org/abs/1803.11496v2","url_pdf":"http://arxiv.org/pdf/1803.11496v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"predicting-future-instance-segmentation-by","repo_url":"https://github.com/facebookresearch/instpred","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"video-forecasting","task_name":"Video Forecasting"},{"task_slug":"video-prediction","task_name":"Video Prediction"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"mask-r-cnn","method_name":"Mask R-CNN"},{"method_slug":"rpn","method_name":"RPN"},{"method_slug":"roi-align","method_name":"RoIAlign"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1803.11496","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}