{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/soc-semantic-assisted-object-cluster-for","title":"SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation","arxiv_id":"2305.17011","date":"2023-05-26","proceeding":"NeurIPS 2023 11","authors":["Zhuoyan Luo","Yicheng Xiao","Yong liu","Shuyan Li","Yitong Wang","Yansong Tang","Xiu Li","Yujiu Yang"],"abstract":"This paper studies referring video object segmentation (RVOS) by boosting video-level visual-linguistic alignment. Recent approaches model the RVOS task as a sequence prediction problem and perform multi-modal interaction as well as segmentation for each frame separately. However, the lack of a global view of video content leads to difficulties in effectively utilizing inter-frame relationships and understanding textual descriptions of object temporal variations. To address this issue, we propose Semantic-assisted Object Cluster (SOC), which aggregates video content and textual guidance for unified temporal modeling and cross-modal alignment. By associating a group of frame-level object embeddings with language tokens, SOC facilitates joint space learning across modalities and time steps. Moreover, we present multi-modal contrastive supervision to help construct well-aligned joint space at the video level. We conduct extensive experiments on popular RVOS benchmarks, and our method outperforms state-of-the-art competitors on all benchmarks by a remarkable margin. Besides, the emphasis on temporal coherence enhances the segmentation stability and adaptability of our method in processing text expressions with temporal variations. Code will be available.","url_abs":"https://arxiv.org/abs/2305.17011v1","url_pdf":"https://arxiv.org/pdf/2305.17011v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"soc-semantic-assisted-object-cluster-for","repo_url":"https://github.com/RobertLuo1/NeurIPS2023_SOC","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"referring-expression-segmentation","task_name":"Referring Expression Segmentation"},{"task_slug":"referring-video-object-segmentation","task_name":"Referring Video Object Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"video-object-segmentation","task_name":"Video Object Segmentation"},{"task_slug":"video-semantic-segmentation","task_name":"Video Semantic Segmentation"},{"task_slug":"cross-modal-alignment","task_name":"cross-modal alignment"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/referring-expression-segmentation-on-a2d","task":"Referring Expression Segmentation","dataset":"A2D Sentences","model":"SOC (Video-Swin-B)","rank_in_archive_order":2,"of":27,"metrics":{"AP":"0.573","IoU mean":"0.725","IoU overall":"0.807","Precision@0.5":"0.851","Precision@0.6":"0.827","Precision@0.7":"0.765","Precision@0.8":"0.607","Precision@0.9":"0.252"},"uses_additional_data":true},{"leaderboard":"/sota/referring-expression-segmentation-on-a2d","task":"Referring Expression Segmentation","dataset":"A2D Sentences","model":"SOC (Video-Swin-T)","rank_in_archive_order":4,"of":27,"metrics":{"AP":"0.504","IoU mean":"0.669","IoU overall":"0.747","Precision@0.5":"0.79","Precision@0.6":"0.756","Precision@0.7":"0.687","Precision@0.8":"0.535","Precision@0.9":"0.195"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-j-hmdb","task":"Referring Expression Segmentation","dataset":"J-HMDB","model":"SOC (Video-Swin-B)","rank_in_archive_order":2,"of":21,"metrics":{"AP":"0.446","IoU mean":"0.723","IoU overall":"0.736","Precision@0.5":"0.969","Precision@0.6":"0.914","Precision@0.7":"0.711","Precision@0.8":"0.213","Precision@0.9":"0.001"},"uses_additional_data":true},{"leaderboard":"/sota/referring-expression-segmentation-on-j-hmdb","task":"Referring Expression Segmentation","dataset":"J-HMDB","model":"SOC (Video-Swin-T)","rank_in_archive_order":4,"of":21,"metrics":{"AP":"0.397","IoU mean":"0.701","IoU overall":"0.707","Precision@0.5":"0.947","Precision@0.6":"0.864","Precision@0.7":"0.627","Precision@0.8":"0.179","Precision@0.9":"0.001"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-refer-1","task":"Referring Expression Segmentation","dataset":"Refer-YouTube-VOS (2021 public validation)","model":"SOC (Joint training, Video-Swin-B)","rank_in_archive_order":9,"of":33,"metrics":{"F":"69.3","J":"65.3","J&F":"67.3±0.5"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-refer-1","task":"Referring Expression Segmentation","dataset":"Refer-YouTube-VOS (2021 public validation)","model":"SOC (Video-Swin-T)","rank_in_archive_order":23,"of":33,"metrics":{"F":"60.5","J":"57.8","J&F":"59.2"},"uses_additional_data":false},{"leaderboard":"/sota/referring-video-object-segmentation-on-long","task":"Referring Video Object Segmentation","dataset":"Long-RVOS","model":"SOC","rank_in_archive_order":6,"of":7,"metrics":{"J&F":"34.9","tIoU":"68.1","vIoU":"28.6"},"uses_additional_data":false},{"leaderboard":"/sota/referring-video-object-segmentation-on-ref","task":"Referring Video Object Segmentation","dataset":"Ref-DAVIS17","model":"SOC","rank_in_archive_order":4,"of":11,"metrics":{"F":"69.1","J":"62.5","J&F":"65.8"},"uses_additional_data":false},{"leaderboard":"/sota/referring-video-object-segmentation-on-refer","task":"Referring Video Object Segmentation","dataset":"Refer-YouTube-VOS","model":"SOC","rank_in_archive_order":6,"of":18,"metrics":{"F":"67.9","J":"64.1","J&F":"66.0"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2305.17011","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}