{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/video-object-segmentation-with-language","title":"Video Object Segmentation with Language Referring Expressions","arxiv_id":"1803.08006","date":"2018-03-21","proceeding":null,"authors":["Anna Khoreva","Anna Rohrbach","Bernt Schiele"],"abstract":"Most state-of-the-art semi-supervised video object segmentation methods rely\non a pixel-accurate mask of a target object provided for the first frame of a\nvideo. However, obtaining a detailed segmentation mask is expensive and\ntime-consuming. In this work we explore an alternative way of identifying a\ntarget object, namely by employing language referring expressions. Besides\nbeing a more practical and natural way of pointing out a target object, using\nlanguage specifications can help to avoid drift as well as make the system more\nrobust to complex dynamics and appearance variations. Leveraging recent\nadvances of language grounding models designed for images, we propose an\napproach to extend them to video data, ensuring temporally coherent\npredictions. To evaluate our method we augment the popular video object\nsegmentation benchmarks, DAVIS'16 and DAVIS'17 with language descriptions of\ntarget objects. We show that our language-supervised approach performs on par\nwith the methods which have access to a pixel-level mask of the target object\non DAVIS'16 and is competitive to methods using scribbles on the challenging\nDAVIS'17 dataset.","url_abs":"http://arxiv.org/abs/1803.08006v3","url_pdf":"http://arxiv.org/pdf/1803.08006v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"referring-expression-segmentation","task_name":"Referring Expression Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"semi-supervised-video-object-segmentation","task_name":"Semi-Supervised Video Object Segmentation"},{"task_slug":"video-object-segmentation","task_name":"Video Object Segmentation"},{"task_slug":"video-semantic-segmentation","task_name":"Video Semantic Segmentation"}],"methods":[],"datasets_introduced":[{"slug":"referring-expressions-for-davis-2016-2017","name":"Referring Expressions for DAVIS 2016 & 2017","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/referring-expression-segmentation-on-davis","task":"Referring Expression Segmentation","dataset":"DAVIS 2017 (val)","model":"Khoreva et al.","rank_in_archive_order":16,"of":18,"metrics":{"J&F 1st frame":"39.3","J&F Full video":"37.1"},"uses_additional_data":false},{"leaderboard":"/sota/visual-object-tracking-on-davis-2016","task":"Semi-Supervised Video Object Segmentation","dataset":"DAVIS 2016","model":"VOSwL","rank_in_archive_order":54,"of":78,"metrics":{"F-measure (Decay)":"8.6","F-measure (Mean)":"84.2","F-measure (Recall)":"93.9","J&F":"83.65","Jaccard (Decay)":"6.9","Jaccard (Mean)":"83.1","Jaccard (Recall)":"95.7"},"uses_additional_data":false},{"leaderboard":"/sota/visual-object-tracking-on-davis-2017","task":"Semi-Supervised Video Object Segmentation","dataset":"DAVIS 2017 (val)","model":"VOSwL (Language)","rank_in_archive_order":71,"of":81,"metrics":{"J&F":"60.8","Jaccard (Mean)":"58.0"},"uses_additional_data":false},{"leaderboard":"/sota/visual-object-tracking-on-davis-2017","task":"Semi-Supervised Video Object Segmentation","dataset":"DAVIS 2017 (val)","model":"VOSwL","rank_in_archive_order":81,"of":81,"metrics":{"F-measure (Decay)":"24.5","F-measure (Mean)":"63.5","F-measure (Recall)":"70.4","Jaccard (Decay)":"22.4","Jaccard (Recall)":"66.1"},"uses_additional_data":false},{"leaderboard":"/sota/video-object-segmentation-on-davis-2016","task":"Video Object Segmentation","dataset":"DAVIS 2016","model":"VOSwL (Mask+Language)","rank_in_archive_order":21,"of":24,"metrics":{"mIoU":"84.5"},"uses_additional_data":false},{"leaderboard":"/sota/video-object-segmentation-on-davis-2016","task":"Video Object Segmentation","dataset":"DAVIS 2016","model":"VOSwL (Language)","rank_in_archive_order":22,"of":24,"metrics":{"mIoU":"82.8"},"uses_additional_data":false},{"leaderboard":"/sota/video-object-segmentation-on-davis-2017","task":"Video Object Segmentation","dataset":"DAVIS 2017","model":"VOSwL (Mask+Language)","rank_in_archive_order":3,"of":5,"metrics":{"J&F":"62.2","mIoU":"59"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1803.08006","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}