{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/domain-alignment-and-temporal-aggregation-for","title":"Dual Prototype Attention for Unsupervised Video Object Segmentation","arxiv_id":"2211.12036","date":"2022-11-22","proceeding":"CVPR 2024 1","authors":["Suhwan Cho","Minhyeok Lee","Seunghoon Lee","Dogyoon Lee","Heeseung Choi","Ig-Jae Kim","Sangyoun Lee"],"abstract":"Unsupervised video object segmentation (VOS) aims to detect and segment the most salient object in videos. The primary techniques used in unsupervised VOS are 1) the collaboration of appearance and motion information; and 2) temporal fusion between different frames. This paper proposes two novel prototype-based attention mechanisms, inter-modality attention (IMA) and inter-frame attention (IFA), to incorporate these techniques via dense propagation across different modalities and frames. IMA densely integrates context information from different modalities based on a mutual refinement. IFA injects global context of a video to the query frame, enabling a full utilization of useful properties from multiple frames. Experimental results on public benchmark datasets demonstrate that our proposed approach outperforms all existing methods by a substantial margin. The proposed two components are also thoroughly validated via ablative study.","url_abs":"https://arxiv.org/abs/2211.12036v3","url_pdf":"https://arxiv.org/pdf/2211.12036v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"domain-alignment-and-temporal-aggregation-for","repo_url":"https://github.com/hydragon516/dpa","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"unsupervised-video-object-segmentation","task_name":"Unsupervised Video Object Segmentation"},{"task_slug":"video-object-segmentation","task_name":"Video Object Segmentation"},{"task_slug":"video-semantic-segmentation","task_name":"Video Semantic Segmentation"}],"methods":[{"method_slug":"tam","method_name":"TAM"},{"method_slug":"vos","method_name":"VOS"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/unsupervised-video-object-segmentation-on-10","task":"Unsupervised Video Object Segmentation","dataset":"DAVIS 2016 val","model":"DPA","rank_in_archive_order":4,"of":25,"metrics":{"F":"88.4","G":"87.6","J":"86.8"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-video-object-segmentation-on-11","task":"Unsupervised Video Object Segmentation","dataset":"FBMS test","model":"DPA","rank_in_archive_order":2,"of":15,"metrics":{"J":"83.4"},"uses_additional_data":false},{"leaderboard":"/sota/unsupervised-video-object-segmentation-on-12","task":"Unsupervised Video Object Segmentation","dataset":"YouTube-Objects","model":"DPA","rank_in_archive_order":3,"of":16,"metrics":{"J":"73.7"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}