Papers › Cross-Modal Progressive Comprehension for Referring Segmentation

Cross-Modal Progressive Comprehension for Referring Segmentation

15 May 2021arXiv:2105.07175archive 2025-07-28

Si Liu, Tianrui Hui, Shaofei Huang, Yunchao Wei, Bo Li, Guanbin Li

Given a natural language expression and an image/video, the goal of referring segmentation is to produce the pixel-level masks of the entities described by the subject of the expression. Previous approaches tackle this problem by implicit feature interaction and fusion between visual and linguistic modalities in a one-stage manner. However, human tends to solve the referring problem in a progressive manner based on informative words in the expression, i.e., first roughly locating candidate entities and then distinguishing the target one. In this paper, we propose a Cross-Modal Progressive Comprehension (CMPC) scheme to effectively mimic human behaviors and implement it as a CMPC-I (Image) module and a CMPC-V (Video) module to improve referring image and video segmentation models. For image data, our CMPC-I module first employs entity and attribute words to perceive all the related entities that might be considered by the expression. Then, the relational words are adopted to highlight the target entity as well as suppress other irrelevant ones by spatial graph reasoning. For video data, our CMPC-V module further exploits action words based on CMPC-I to highlight the correct entity matched with the action cues by temporal graph reasoning. In addition to the CMPC, we also introduce a simple yet effective Text-Guided Feature Exchange (TGFE) module to integrate the reasoned multimodal features corresponding to different levels in the visual backbone under the guidance of textual information. In this way, multi-level features can communicate with each other and be mutually refined based on the textual context. Combining CMPC-I or CMPC-V with TGFE can form our image or video version referring segmentation frameworks and our frameworks achieve new state-of-the-art performances on four referring image segmentation benchmarks and three referring video segmentation benchmarks respectively.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

spyflying/CMPC-Refseg officialmentioned in papertfMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

AttributeImage SegmentationReferring Expression SegmentationSegmentationSemantic SegmentationVideo SegmentationVideo Semantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Referring Expression Segmentation A2D Sentences CMPC-V (I3D) AP 0.404 #11 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMPC-V (I3D) IoU mean 0.573 #11 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMPC-V (I3D) IoU overall 0.653 #11 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMPC-V (I3D) Precision@0.5 0.655 #11 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMPC-V (I3D) Precision@0.6 0.592 #11 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMPC-V (I3D) Precision@0.7 0.506 #11 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMPC-V (I3D) Precision@0.8 0.342 #11 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMPC-V (I3D) Precision@0.9 0.098 #11 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMPC-V (R2D) AP 0.351 #15 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMPC-V (R2D) IoU mean 0.515 #15 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMPC-V (R2D) IoU overall 0.649 #15 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMPC-V (R2D) Precision@0.5 0.590 #15 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMPC-V (R2D) Precision@0.6 0.527 #15 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMPC-V (R2D) Precision@0.7 0.434 #15 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMPC-V (R2D) Precision@0.8 0.284 #15 of 27 Archive leaderboard report
Referring Expression Segmentation A2D Sentences CMPC-V (R2D) Precision@0.9 0.068 #15 of 27 Archive leaderboard report
Referring Expression Segmentation J-HMDB CMPC-V AP 0.342 #7 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB CMPC-V IoU mean 0.617 #7 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB CMPC-V IoU overall 0.616 #7 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB CMPC-V Precision@0.5 0.813 #7 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB CMPC-V Precision@0.6 0.657 #7 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB CMPC-V Precision@0.7 0.371 #7 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB CMPC-V Precision@0.8 0.07 #7 of 21 Archive leaderboard report
Referring Expression Segmentation J-HMDB CMPC-V Precision@0.9 0.000 #7 of 21 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections