Papers › SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation

SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation

26 May 2023NeurIPS 2023 11arXiv:2305.17011archive 2025-07-28

Zhuoyan Luo, Yicheng Xiao, Yong liu, Shuyan Li, Yitong Wang, Yansong Tang, Xiu Li, Yujiu Yang

This paper studies referring video object segmentation (RVOS) by boosting video-level visual-linguistic alignment. Recent approaches model the RVOS task as a sequence prediction problem and perform multi-modal interaction as well as segmentation for each frame separately. However, the lack of a global view of video content leads to difficulties in effectively utilizing inter-frame relationships and understanding textual descriptions of object temporal variations. To address this issue, we propose Semantic-assisted Object Cluster (SOC), which aggregates video content and textual guidance for unified temporal modeling and cross-modal alignment. By associating a group of frame-level object embeddings with language tokens, SOC facilitates joint space learning across modalities and time steps. Moreover, we present multi-modal contrastive supervision to help construct well-aligned joint space at the video level. We conduct extensive experiments on popular RVOS benchmarks, and our method outperforms state-of-the-art competitors on all benchmarks by a remarkable margin. Besides, the emphasis on temporal coherence enhances the segmentation stability and adaptability of our method in processing text expressions with temporal variations. Code will be available.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

RobertLuo1/NeurIPS2023_SOC officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

ObjectReferring Expression SegmentationReferring Video Object SegmentationSegmentationSemantic SegmentationVideo Object SegmentationVideo Semantic Segmentationcross-modal alignment

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Referring Expression Segmentation A2D Sentences SOC (Video-Swin-B) AP 0.573 #2 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences SOC (Video-Swin-B) IoU mean 0.725 #2 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences SOC (Video-Swin-B) IoU overall 0.807 #2 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences SOC (Video-Swin-B) Precision@0.5 0.851 #2 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences SOC (Video-Swin-B) Precision@0.6 0.827 #2 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences SOC (Video-Swin-B) Precision@0.7 0.765 #2 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences SOC (Video-Swin-B) Precision@0.8 0.607 #2 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences SOC (Video-Swin-B) Precision@0.9 0.252 #2 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences SOC (Video-Swin-T) AP 0.504 #4 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences SOC (Video-Swin-T) IoU mean 0.669 #4 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences SOC (Video-Swin-T) IoU overall 0.747 #4 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences SOC (Video-Swin-T) Precision@0.5 0.79 #4 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences SOC (Video-Swin-T) Precision@0.6 0.756 #4 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences SOC (Video-Swin-T) Precision@0.7 0.687 #4 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences SOC (Video-Swin-T) Precision@0.8 0.535 #4 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences SOC (Video-Swin-T) Precision@0.9 0.195 #4 of 27 Archive leaderboard report
Referring Expression Segmentation J-HMDB SOC (Video-Swin-B) AP 0.446 #2 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB SOC (Video-Swin-B) IoU mean 0.723 #2 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB SOC (Video-Swin-B) IoU overall 0.736 #2 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB SOC (Video-Swin-B) Precision@0.5 0.969 #2 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB SOC (Video-Swin-B) Precision@0.6 0.914 #2 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB SOC (Video-Swin-B) Precision@0.7 0.711 #2 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB SOC (Video-Swin-B) Precision@0.8 0.213 #2 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB SOC (Video-Swin-B) Precision@0.9 0.001 #2 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB SOC (Video-Swin-T) AP 0.397 #4 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB SOC (Video-Swin-T) IoU mean 0.701 #4 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB SOC (Video-Swin-T) IoU overall 0.707 #4 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB SOC (Video-Swin-T) Precision@0.5 0.947 #4 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB SOC (Video-Swin-T) Precision@0.6 0.864 #4 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB SOC (Video-Swin-T) Precision@0.7 0.627 #4 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB SOC (Video-Swin-T) Precision@0.8 0.179 #4 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB SOC (Video-Swin-T) Precision@0.9 0.001 #4 of 21 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS (2021 public validation) SOC (Joint training, Video-Swin-B) F 69.3 #9 of 33 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS (2021 public validation) SOC (Joint training, Video-Swin-B) J 65.3 #9 of 33 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS (2021 public validation) SOC (Joint training, Video-Swin-B) J&F 67.3±0.5 #9 of 33 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS (2021 public validation) SOC (Video-Swin-T) F 60.5 #23 of 33 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS (2021 public validation) SOC (Video-Swin-T) J 57.8 #23 of 33 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS (2021 public validation) SOC (Video-Swin-T) J&F 59.2 #23 of 33 Archive leaderboard report
Referring Video Object Segmentation Long-RVOS SOC J&F 34.9 #6 of 7 Archive leaderboard report
Referring Video Object Segmentation Long-RVOS SOC tIoU 68.1 #6 of 7 Archive leaderboard report
Referring Video Object Segmentation Long-RVOS SOC vIoU 28.6 #6 of 7 Archive leaderboard report
Referring Video Object Segmentation Ref-DAVIS17 SOC F 69.1 #4 of 11 Archive leaderboard report
Referring Video Object Segmentation Ref-DAVIS17 SOC J 62.5 #4 of 11 Archive leaderboard report
Referring Video Object Segmentation Ref-DAVIS17 SOC J&F 65.8 #4 of 11 Archive leaderboard report
Referring Video Object Segmentation Refer-YouTube-VOS SOC F 67.9 #6 of 18 Archive leaderboard report
Referring Video Object Segmentation Refer-YouTube-VOS SOC J 64.1 #6 of 18 Archive leaderboard report
Referring Video Object Segmentation Refer-YouTube-VOS SOC J&F 66.0 #6 of 18 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections