{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/semantic-video-segmentation-by-gated","title":"Semantic Video Segmentation by Gated Recurrent Flow Propagation","arxiv_id":"1612.08871","date":"2016-12-28","proceeding":"CVPR 2018 6","authors":["David Nilsson","Cristian Sminchisescu"],"abstract":"Semantic video segmentation is challenging due to the sheer amount of data\nthat needs to be processed and labeled in order to construct accurate models.\nIn this paper we present a deep, end-to-end trainable methodology to video\nsegmentation that is capable of leveraging information present in unlabeled\ndata in order to improve semantic estimates. Our model combines a convolutional\narchitecture and a spatio-temporal transformer recurrent layer that are able to\ntemporally propagate labeling information by means of optical flow, adaptively\ngated based on its locally estimated uncertainty. The flow, the recognition and\nthe gated temporal propagation modules can be trained jointly, end-to-end. The\ntemporal, gated recurrent flow propagation component of our model can be\nplugged into any static semantic segmentation architecture and turn it into a\nweakly supervised video processing one. Our extensive experiments in the\nchallenging CityScapes and Camvid datasets, and based on multiple deep\narchitectures, indicate that the resulting model can leverage unlabeled\ntemporal frames, next to a labeled one, in order to improve both the video\nsegmentation accuracy and the consistency of its temporal labeling, at no\nadditional annotation cost and with little extra computation.","url_abs":"http://arxiv.org/abs/1612.08871v2","url_pdf":"http://arxiv.org/pdf/1612.08871v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"optical-flow-estimation","task_name":"Optical Flow Estimation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"video-segmentation","task_name":"Video Segmentation"},{"task_slug":"video-semantic-segmentation","task_name":"Video Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-semantic-segmentation-on-camvid","task":"Video Semantic Segmentation","dataset":"CamVid","model":"GRFP","rank_in_archive_order":6,"of":6,"metrics":{"Mean IoU":"67.1"},"uses_additional_data":false},{"leaderboard":"/sota/video-semantic-segmentation-on-cityscapes-val","task":"Video Semantic Segmentation","dataset":"Cityscapes val","model":"GRFP [15]","rank_in_archive_order":7,"of":9,"metrics":{"mIoU":"73.6"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1612.08871","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}