{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/video-object-segmentation-without-temporal","title":"Video Object Segmentation Without Temporal Information","arxiv_id":"1709.06031","date":"2017-09-18","proceeding":null,"authors":["Kevis-Kokitsi Maninis","Sergi Caelles","Yu-Hua Chen","Jordi Pont-Tuset","Laura Leal-Taixé","Daniel Cremers","Luc van Gool"],"abstract":"Video Object Segmentation, and video processing in general, has been\nhistorically dominated by methods that rely on the temporal consistency and\nredundancy in consecutive video frames. When the temporal smoothness is\nsuddenly broken, such as when an object is occluded, or some frames are missing\nin a sequence, the result of these methods can deteriorate significantly or\nthey may not even produce any result at all. This paper explores the orthogonal\napproach of processing each frame independently, i.e disregarding the temporal\ninformation. In particular, it tackles the task of semi-supervised video object\nsegmentation: the separation of an object from the background in a video, given\nits mask in the first frame. We present Semantic One-Shot Video Object\nSegmentation (OSVOS-S), based on a fully-convolutional neural network\narchitecture that is able to successively transfer generic semantic\ninformation, learned on ImageNet, to the task of foreground segmentation, and\nfinally to learning the appearance of a single annotated object of the test\nsequence (hence one shot). We show that instance level semantic information,\nwhen combined effectively, can dramatically improve the results of our previous\nmethod, OSVOS. We perform experiments on two recent video segmentation\ndatabases, which show that OSVOS-S is both the fastest and most accurate method\nin the state of the art.","url_abs":"http://arxiv.org/abs/1709.06031v2","url_pdf":"http://arxiv.org/pdf/1709.06031v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"foreground-segmentation","task_name":"Foreground Segmentation"},{"task_slug":"object","task_name":"Object"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"semi-supervised-video-object-segmentation","task_name":"Semi-Supervised Video Object Segmentation"},{"task_slug":"video-object-segmentation","task_name":"Video Object Segmentation"},{"task_slug":"video-segmentation","task_name":"Video Segmentation"},{"task_slug":"video-semantic-segmentation","task_name":"Video Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-object-tracking-on-davis-2016","task":"Semi-Supervised Video Object Segmentation","dataset":"DAVIS 2016","model":"OSVOS-S","rank_in_archive_order":47,"of":78,"metrics":{"F-measure (Decay)":"8.2","F-measure (Mean)":"87.5","F-measure (Recall)":"95.9","J&F":"86.55","Jaccard (Decay)":"5.5","Jaccard (Mean)":"85.6","Jaccard (Recall)":"96.8"},"uses_additional_data":false},{"leaderboard":"/sota/semi-supervised-video-object-segmentation-on-1","task":"Semi-Supervised Video Object Segmentation","dataset":"DAVIS 2017 (test-dev)","model":"OSVOS-S","rank_in_archive_order":48,"of":59,"metrics":{"F-measure (Decay)":"21.9","F-measure (Mean)":"62.1","F-measure (Recall)":"70.5","J&F":"57.5","Jaccard (Decay)":"24.1","Jaccard (Mean)":"52.9","Jaccard (Recall)":"60.2"},"uses_additional_data":false},{"leaderboard":"/sota/visual-object-tracking-on-davis-2017","task":"Semi-Supervised Video Object Segmentation","dataset":"DAVIS 2017 (val)","model":"OSVOS-S","rank_in_archive_order":63,"of":81,"metrics":{"F-measure (Decay)":"18.5","F-measure (Mean)":"71.3","F-measure (Recall)":"80.7","J&F":"68","Jaccard (Decay)":"15.1","Jaccard (Mean)":"64.7","Jaccard (Recall)":"74.2"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1709.06031","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}