Papers › Deeply Interleaved Two-Stream Encoder for Referring Video Segmentation

Deeply Interleaved Two-Stream Encoder for Referring Video Segmentation

30 Mar 2022arXiv:2203.15969archive 2025-07-28

Guang Feng, Lihe Zhang, Zhiwei Hu, Huchuan Lu

Referring video segmentation aims to segment the corresponding video object described by the language expression. To address this task, we first design a two-stream encoder to extract CNN-based visual features and transformer-based linguistic features hierarchically, and a vision-language mutual guidance (VLMG) module is inserted into the encoder multiple times to promote the hierarchical and progressive fusion of multi-modal features. Compared with the existing multi-modal fusion methods, this two-stream encoder takes into account the multi-granularity linguistic context, and realizes the deep interleaving between modalities with the help of VLGM. In order to promote the temporal alignment between frames, we further propose a language-guided multi-scale dynamic filtering (LMDF) module to strengthen the temporal coherence, which uses the language-guided spatial-temporal features to generate a set of position-specific dynamic filters to more flexibly and effectively update the feature of current frame. Extensive experiments on four datasets verify the effectiveness of the proposed model.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Referring Expression SegmentationVideo SegmentationVideo Semantic SegmentationVocal Bursts Valence Prediction

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Referring Expression Segmentation A2D Sentences VLIDE AP 0.469 #6 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences VLIDE IoU mean 0.598 #6 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences VLIDE IoU overall 0.714 #6 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences VLIDE Precision@0.5 0.702 #6 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences VLIDE Precision@0.6 0.663 #6 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences VLIDE Precision@0.7 0.585 #6 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences VLIDE Precision@0.8 0.428 #6 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences VLIDE Precision@0.9 0.151 #6 of 27 Archive leaderboard report
Referring Expression Segmentation J-HMDB VLIDE AP 0.441 #3 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB VLIDE IoU mean 0.666 #3 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB VLIDE IoU overall 0.68 #3 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB VLIDE Precision@0.5 0.874 #3 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB VLIDE Precision@0.6 0.791 #3 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB VLIDE Precision@0.7 0.586 #3 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB VLIDE Precision@0.8 0.182 #3 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB VLIDE Precision@0.9 0.30 #3 of 21 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS (2021 public validation) VLIDE F 50.67 #31 of 33 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS (2021 public validation) VLIDE J 48.44 #31 of 33 Archive leaderboard report
Referring Expression Segmentation Refer-YouTube-VOS (2021 public validation) VLIDE J&F 49.56 #31 of 33 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections