Papers › Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation

Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation

1 Aug 2024arXiv:2408.00744archive 2025-07-28

Siyu Jiao, Hongguang Zhu, Jiannan Huang, Yao Zhao, Yunchao Wei, Humphrey Shi

Pre-trained vision-language models, e.g. CLIP, have been increasingly used to address the challenging Open-Vocabulary Segmentation (OVS) task, benefiting from their well-aligned vision-text embedding space. Typical solutions involve either freezing CLIP during training to unilaterally maintain its zero-shot capability, or fine-tuning CLIP vision encoder to achieve perceptual sensitivity to local regions. However, few of them incorporate vision-text collaborative optimization. Based on this, we propose the Content-Dependent Transfer to adaptively enhance each text embedding by interacting with the input image, which presents a parameter-efficient way to optimize the text representation. Besides, we additionally introduce a Representation Compensation strategy, reviewing the original CLIP-V representation as compensation to maintain the zero-shot capability of CLIP. In this way, the vision and text representation of CLIP are optimized collaboratively, enhancing the alignment of the vision-text feature space. To the best of our knowledge, we are the first to establish the collaborative vision-text optimizing mechanism within the OVS field. Extensive experiments demonstrate our method achieves superior performance on popular OVS benchmarks. In open-vocabulary semantic segmentation, our method outperforms the previous state-of-the-art approaches by +0.5, +2.3, +3.4, +0.4 and +1.1 mIoU, respectively on A-847, A-150, PC-459, PC-59 and PAS-20. Furthermore, in a panoptic setting on ADE20K, we achieve the performance of 27.1 PQ, 73.5 SQ, and 32.9 RQ. Code will be available at https://github.com/jiaosiyu1999/MAFT-Plus.git .

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

jiaosiyu1999/MAFT-Plus officialmentioned in papermentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Open Vocabulary Panoptic SegmentationOpen Vocabulary Semantic SegmentationSemantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Open Vocabulary Panoptic Segmentation ADE20K MAFT+ PQ 27.1 #4 of 10 Archive leaderboard report
Open Vocabulary Semantic Segmentation ADE20K-150 MAFT+ mIoU 36.1 #6 of 23 Archive leaderboard report
Open Vocabulary Semantic Segmentation ADE20K-847 MAFT+ mIoU 15.1 #6 of 19 Archive leaderboard report
Open Vocabulary Semantic Segmentation PASCAL Context-459 MAFT+ mIoU 21.6 #8 of 15 Archive leaderboard report
Open Vocabulary Semantic Segmentation PASCAL Context-59 MAFT+ mIoU 59.4 #10 of 24 Archive leaderboard report
Open Vocabulary Semantic Segmentation PascalVOC-20 MAFT+ mIoU 96.5 #6 of 20 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

CLIP

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections