{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/see-more-know-more-unsupervised-video-object-1","title":"See More, Know More: Unsupervised Video Object Segmentation with Co-Attention Siamese Networks","arxiv_id":"2001.06810","date":"2020-01-19","proceeding":"CVPR 2019 6","authors":["Xiankai Lu","Wenguan Wang","Chao Ma","Jianbing Shen","Ling Shao","Fatih Porikli"],"abstract":"We introduce a novel network, called CO-attention Siamese Network (COSNet), to address the unsupervised video object segmentation task from a holistic view. We emphasize the importance of inherent correlation among video frames and incorporate a global co-attention mechanism to improve further the state-of-the-art deep learning based solutions that primarily focus on learning discriminative foreground representations over appearance and motion in short-term temporal segments. The co-attention layers in our network provide efficient and competent stages for capturing global correlations and scene context by jointly computing and appending co-attention responses into a joint feature space. We train COSNet with pairs of video frames, which naturally augments training data and allows increased learning capacity. During the segmentation stage, the co-attention model encodes useful information by processing multiple reference frames together, which is leveraged to infer the frequently reappearing and salient foreground objects better. We propose a unified and end-to-end trainable framework where different co-attention variants can be derived for mining the rich context within videos. Our extensive experiments over three large benchmarks manifest that COSNet outperforms the current alternatives by a large margin.","url_abs":"https://arxiv.org/abs/2001.06810v1","url_pdf":"https://arxiv.org/pdf/2001.06810v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"see-more-know-more-unsupervised-video-object-1","repo_url":"https://github.com/carrierlxk/COSNet","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"unsupervised-video-object-segmentation","task_name":"Unsupervised Video Object Segmentation"},{"task_slug":"video-object-segmentation","task_name":"Video Object Segmentation"},{"task_slug":"video-polyp-segmentation","task_name":"Video Polyp Segmentation"},{"task_slug":"video-semantic-segmentation","task_name":"Video Semantic Segmentation"}],"methods":[{"method_slug":"siamese-network","method_name":"Siamese Network"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/unsupervised-video-object-segmentation-on-10","task":"Unsupervised Video Object Segmentation","dataset":"DAVIS 2016 val","model":"COSNet","rank_in_archive_order":22,"of":25,"metrics":{"F":"79.4","G":"80.0","J":"80.5"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-video-object-segmentation-on-11","task":"Unsupervised Video Object Segmentation","dataset":"FBMS test","model":"COSNet","rank_in_archive_order":14,"of":15,"metrics":{"J":"75.6"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-video-object-segmentation-on-12","task":"Unsupervised Video Object Segmentation","dataset":"YouTube-Objects","model":"COSNet","rank_in_archive_order":11,"of":16,"metrics":{"J":"70.5"},"uses_additional_data":false},{"leaderboard":"/sota/video-polyp-segmentation-on-sun-seg-easy","task":"Video Polyp Segmentation","dataset":"SUN-SEG-Easy (Unseen)","model":"COSNet","rank_in_archive_order":13,"of":18,"metrics":{"Dice":"0.596","S measure":"0.654","Sensitivity":"0.359","mean E-measure":"0.600","mean F-measure":"0.496","weighted F-measure":"0.431"},"uses_additional_data":false},{"leaderboard":"/sota/video-polyp-segmentation-on-sun-seg-hard","task":"Video Polyp Segmentation","dataset":"SUN-SEG-Hard (Unseen)","model":"COSNet","rank_in_archive_order":11,"of":18,"metrics":{"Dice":"0.606","S-Measure":"0.670","Sensitivity":"0.380","mean E-measure":"0.627","mean F-measure":"0.506","weighted F-measure":"0.443"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2001.06810","atlas_url":"https://app.syntology.ai/?focus=2001.06810","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}