Papers › Visual and Textual Prior Guided Mask Assemble for Few-Shot Segmentation and Beyond

Visual and Textual Prior Guided Mask Assemble for Few-Shot Segmentation and Beyond

15 Aug 2023arXiv:2308.07539archive 2025-07-28

Chen Shuai, Meng Fanman, Zhang Runtong, Qiu Heqian, Li Hongliang, Wu Qingbo, Xu Linfeng

Few-shot segmentation (FSS) aims to segment the novel classes with a few annotated images. Due to CLIP's advantages of aligning visual and textual information, the integration of CLIP can enhance the generalization ability of FSS model. However, even with the CLIP model, the existing CLIP-based FSS methods are still subject to the biased prediction towards base classes, which is caused by the class-specific feature level interactions. To solve this issue, we propose a visual and textual Prior Guided Mask Assemble Network (PGMA-Net). It employs a class-agnostic mask assembly process to alleviate the bias, and formulates diverse tasks into a unified manner by assembling the prior through affinity. Specifically, the class-relevant textual and visual features are first transformed to class-agnostic prior in the form of probability map. Then, a Prior-Guided Mask Assemble Module (PGMAM) including multiple General Assemble Units (GAUs) is introduced. It considers diverse and plug-and-play interactions, such as visual-textual, inter- and intra-image, training-free, and high-order ones. Lastly, to ensure the class-agnostic ability, a Hierarchical Decoder with Channel-Drop Mechanism (HDCDM) is proposed to flexibly exploit the assembled masks and low-level features, without relying on any class-specific information. It achieves new state-of-the-art results in the FSS task, with mIoU of $77.6$ on PASCAL-5ⁱ and $59.4$ on COCO-20ⁱ in 1-shot scenario. Beyond this, we show that without extra re-training, the proposed PGMA-Net can solve bbox-level and cross-domain FSS, co-segmentation, zero-shot segmentation (ZSS) tasks, leading an any-shot segmentation framework.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Few-Shot Semantic SegmentationSegmentationZero Shot Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Few-Shot Semantic Segmentation COCO-20i (1-shot) PGMA-Net (ResNet-101) FB-IoU 78.5 #1 of 85 Archive leaderboard report
Few-Shot Semantic Segmentation COCO-20i (1-shot) PGMA-Net (ResNet-101) Mean IoU 59.4 #1 of 85 Archive leaderboard report
Few-Shot Semantic Segmentation COCO-20i (1-shot) PGMA-Net (ResNet-50) FB-IoU 75.8 #4 of 85 Archive leaderboard report
Few-Shot Semantic Segmentation COCO-20i (1-shot) PGMA-Net (ResNet-50) Mean IoU 54.3 #4 of 85 Archive leaderboard report
Few-Shot Semantic Segmentation COCO-20i (5-shot) PGMA-Net (ResNet-101) FB-IoU 79.4 #3 of 81 Archive leaderboard report
Few-Shot Semantic Segmentation COCO-20i (5-shot) PGMA-Net (ResNet-101) Mean IoU 61.8 #3 of 81 Archive leaderboard report
Few-Shot Semantic Segmentation COCO-20i (5-shot) PGMA-Net (ResNet-50) FB-IoU 76.7 #11 of 81 Archive leaderboard report
Few-Shot Semantic Segmentation COCO-20i (5-shot) PGMA-Net (ResNet-50) Mean IoU 57.1 #11 of 81 Archive leaderboard report
Few-Shot Semantic Segmentation PASCAL-5i (1-Shot) PGMA-Net (ResNet-101) FB-IoU 86.2 #2 of 105 Archive leaderboard report
Few-Shot Semantic Segmentation PASCAL-5i (1-Shot) PGMA-Net (ResNet-101) Mean IoU 77.6 #2 of 105 Archive leaderboard report
Few-Shot Semantic Segmentation PASCAL-5i (1-Shot) PGMA-Net (ResNet-50) FB-IoU 83.5 #3 of 105 Archive leaderboard report
Few-Shot Semantic Segmentation PASCAL-5i (1-Shot) PGMA-Net (ResNet-50) Mean IoU 74.1 #3 of 105 Archive leaderboard report
Few-Shot Semantic Segmentation PASCAL-5i (1-Shot) PGMA-Net (ViT-B/16) FB-IoU 82.1 #4 of 105 Archive leaderboard report
Few-Shot Semantic Segmentation PASCAL-5i (1-Shot) PGMA-Net (ViT-B/16) Mean IoU 74.1 #4 of 105 Archive leaderboard report
Few-Shot Semantic Segmentation PASCAL-5i (5-Shot) PGMA-Net (ResNet-101) FB-IoU 86.9 #3 of 96 Archive leaderboard report
Few-Shot Semantic Segmentation PASCAL-5i (5-Shot) PGMA-Net (ResNet-101) Mean IoU 78.6 #3 of 96 Archive leaderboard report
Few-Shot Semantic Segmentation PASCAL-5i (5-Shot) PGMA-Net (ResNet-50) FB-IoU 84.2 #6 of 96 Archive leaderboard report
Few-Shot Semantic Segmentation PASCAL-5i (5-Shot) PGMA-Net (ResNet-50) Mean IoU 75.2 #6 of 96 Archive leaderboard report
Few-Shot Semantic Segmentation PASCAL-5i (5-Shot) PGMA-Net (ViT-B/16) FB-IoU 82.5 #8 of 96 Archive leaderboard report
Few-Shot Semantic Segmentation PASCAL-5i (5-Shot) PGMA-Net (ViT-B/16) Mean IoU 74.6 #8 of 96 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

BASECLIP

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections