{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/online-adaptation-of-convolutional-neural","title":"Online Adaptation of Convolutional Neural Networks for Video Object Segmentation","arxiv_id":"1706.09364","date":"2017-06-28","proceeding":null,"authors":["Paul Voigtlaender","Bastian Leibe"],"abstract":"We tackle the task of semi-supervised video object segmentation, i.e.\nsegmenting the pixels belonging to an object in the video using the ground\ntruth pixel mask for the first frame. We build on the recently introduced\none-shot video object segmentation (OSVOS) approach which uses a pretrained\nnetwork and fine-tunes it on the first frame. While achieving impressive\nperformance, at test time OSVOS uses the fine-tuned network in unchanged form\nand is not able to adapt to large changes in object appearance. To overcome\nthis limitation, we propose Online Adaptive Video Object Segmentation (OnAVOS)\nwhich updates the network online using training examples selected based on the\nconfidence of the network and the spatial configuration. Additionally, we add a\npretraining step based on objectness, which is learned on PASCAL. Our\nexperiments show that both extensions are highly effective and improve the\nstate of the art on DAVIS to an intersection-over-union score of 85.7%.","url_abs":"http://arxiv.org/abs/1706.09364v2","url_pdf":"http://arxiv.org/pdf/1706.09364v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"semi-supervised-video-object-segmentation","task_name":"Semi-Supervised Video Object Segmentation"},{"task_slug":"video-object-segmentation","task_name":"Video Object Segmentation"},{"task_slug":"video-semantic-segmentation","task_name":"Video Semantic Segmentation"},{"task_slug":"visual-object-tracking","task_name":"Visual Object Tracking"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-object-tracking-on-davis-2016","task":"Semi-Supervised Video Object Segmentation","dataset":"DAVIS 2016","model":"OnAVOS","rank_in_archive_order":50,"of":78,"metrics":{"F-measure (Decay)":"5.8","F-measure (Mean)":"84.9","F-measure (Recall)":"89.7","J&F":"85.5","Jaccard (Decay)":"5.2","Jaccard (Mean)":"86.1","Jaccard (Recall)":"96.1"},"uses_additional_data":false},{"leaderboard":"/sota/semi-supervised-video-object-segmentation-on-1","task":"Semi-Supervised Video Object Segmentation","dataset":"DAVIS 2017 (test-dev)","model":"OnAVOS","rank_in_archive_order":52,"of":59,"metrics":{"F-measure (Decay)":"23.4","F-measure (Recall)":"60.3","J&F":"52.8","Jaccard (Decay)":"23.0","Jaccard (Mean)":"49.9","Jaccard (Recall)":"54.3"},"uses_additional_data":false},{"leaderboard":"/sota/visual-object-tracking-on-davis-2017","task":"Semi-Supervised Video Object Segmentation","dataset":"DAVIS 2017 (val)","model":"OnAVOS","rank_in_archive_order":68,"of":81,"metrics":{"F-measure (Decay)":"26.6","F-measure (Mean)":"69.1","F-measure (Recall)":"75.4","J&F":"65.35","Jaccard (Decay)":"27.9","Jaccard (Mean)":"61.6","Jaccard (Recall)":"67.4"},"uses_additional_data":false},{"leaderboard":"/sota/video-object-segmentation-on-youtube","task":"Semi-Supervised Video Object Segmentation","dataset":"YouTube","model":"OnAVOS","rank_in_archive_order":4,"of":5,"metrics":{"mIoU":"0.774"},"uses_additional_data":false},{"leaderboard":"/sota/visual-object-tracking-on-youtube-vos","task":"Visual Object Tracking","dataset":"YouTube-VOS 2018","model":"OnAVOS","rank_in_archive_order":2,"of":9,"metrics":{"F-Measure (Seen)":"62.7","F-Measure (Unseen)":"51.4","Jaccard (Seen)":"60.1","Jaccard (Unseen)":"46.6","O (Average of Measures)":"55.2"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1706.09364","atlas_url":"https://app.syntology.ai/?focus=1706.09364","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}