Papers › Exploring Regional Clues in CLIP for Zero-Shot Semantic Segmentation

Exploring Regional Clues in CLIP for Zero-Shot Semantic Segmentation

1 Jan 2024CVPR 2024 1archive 2025-07-28

Yi Zhang, Meng-Hao Guo, Miao Wang, Shi-Min Hu

CLIP has demonstrated marked progress in visual recognition due to its powerful pre-training on large-scale image-text pairs. However it still remains a critical challenge: how to transfer image-level knowledge into pixel-level understanding tasks such as semantic segmentation. In this paper to solve the mentioned challenge we analyze the gap between the capability of the CLIP model and the requirement of the zero-shot semantic segmentation task. Based on our analysis and observations we propose a novel method for zero-shot semantic segmentation dubbed CLIP-RC (CLIP with Regional Clues) bringing two main insights. On the one hand a region-level bridge is necessary to provide fine-grained semantics. On the other hand overfitting should be mitigated during the training stage. Benefiting from the above discoveries CLIP-RC achieves state-of-the-art performance on various zero-shot semantic segmentation benchmarks including PASCAL VOC PASCAL Context and COCO-Stuff 164K. Code will be available at https://github.com/Jittor/JSeg.

PaperPDFCode

Code

Jittor/JSeg mentioned in paperpytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

SegmentationSemantic SegmentationZero-Shot Semantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Zero-Shot Semantic Segmentation COCO-Stuff CLIP-RC Inductive Setting hIoU 41.2 #2 of 15 Archive leaderboard report
Zero-Shot Semantic Segmentation COCO-Stuff CLIP-RC Transductive Setting hIoU 49.7 #2 of 15 Archive leaderboard report
Zero-Shot Semantic Segmentation PASCAL VOC CLIP-RC Inductive Setting hIoU 88.4 #3 of 13 Archive leaderboard report
Zero-Shot Semantic Segmentation PASCAL VOC CLIP-RC Transductive Setting hIoU 93.0 #3 of 13 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

CLIP

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections