Papers › Multi-Level Representation Learning With Semantic Alignment for Referring Video Object...
Multi-Level Representation Learning With Semantic Alignment for Referring Video Object Segmentation
Dongming Wu, Xingping Dong, Ling Shao, Jianbing Shen
Referring video object segmentation (RVOS) is a challenging language-guided video grounding task, which requires comprehensively understanding the semantic information of both video content and language queries for object prediction. However, existing methods adopt multi-modal fusion at a frame-based spatial granularity. The limitation of visual representation is prone to causing vision-language mismatching and producing poor segmentation results. To address this, we propose a novel multi-level representation learning approach, which explores the inherent structure of the video content to provide a set of discriminative visual embedding, enabling more effective vision-language semantic alignment. Specifically, we embed different visual cues in terms of visual granularity, including multi-frame long-temporal information at video level, intra-frame spatial semantics at frame level, and enhanced object-aware feature prior at object level. With the powerful multi-level visual embedding and carefully-designed dynamic alignment, our model can generate a robust representation for accurate video object segmentation. Extensive experiments on Refer-DAVIS_ 17 and Refer-YouTube-VOS demonstrate that our model achieves superior performance both in segmentation accuracy and inference speed.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Referring Expression Segmentation | Refer-YouTube-VOS (2021 public validation) | MLRLSA | F | 48.43 | #30 of 33 | Archive leaderboard | report |
| Referring Expression Segmentation | Refer-YouTube-VOS (2021 public validation) | MLRLSA | J | 50.96 | #30 of 33 | Archive leaderboard | report |
| Referring Expression Segmentation | Refer-YouTube-VOS (2021 public validation) | MLRLSA | J&F | 49.70 | #30 of 33 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections