Browse State-of-the-Art › Zero Shot Segmentation
Zero Shot Segmentation
70 papers with code · 2 benchmarks · 3 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| Segmentation in the Wild (12 rows) | Grounded HQ-SAM | Segment Anything in High Quality | code | Syntology ran 3 of 17 samples · 14 unverified | Compare |
| ADE20K training-free zero-shot segmentation (5 rows) | COSMOS ViT-B/16 | COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 70 papers with code (134 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
9 Mar 2023 10 repositories listed Syntology ran 2 of 5 samples · 3 unverifiedTo effectively fuse language and vision modalities, we conceptually divide a closed-set detector into three phases and propose a tight fusion solution, which includes a feature enhancer, a language-guided query…
-
18 Dec 2021 6 repositories listed Syntology ran 1 of 2 samples · 1 unverified · 2 pointer-only (licence)After training on an extended version of the PhraseCut dataset, our system generates a binary segmentation map for an image based on a free-text prompt or on an additional image expressing the query.
-
2 Jun 2023 4 repositories listed Syntology ran 3 of 17 samples · 14 unverified · 3 pointer-only (licence)HQ-SAM is only trained on the introduced detaset of 44k masks, which takes only 4 hours on 8 GPUs.
-
9 Mar 2025 3 repositories listedTraditional methods for reasoning segmentation rely on supervised fine-tuning with categorical labels and simple descriptions, limiting its out-of-domain generalization and lacking explicit reasoning processes.
-
17 May 2023 3 repositories listed Syntology ran 3 of 4 samples · 1 unverifiedTested on 14 previously unseen datasets, the One-Prompt Model showcases superior zero-shot segmentation capabilities, outperforming a wide range of related methods.
-
23 Feb 2023 3 repositories listed Syntology ran 1 of 9 samples · 8 unverifiedA side network is attached to a frozen CLIP model with two branches: one for predicting mask proposals, and the other for predicting attention bias which is applied in the CLIP model to recognize the class of masks.
-
28 Sep 2024 2 repositories listedRecently, the introduction of foundation models like CLIP and Segment-Anything-Model (SAM), with robust cross-domain representations, has paved the way for interactive and universal image segmentation.
-
30 Sep 2023 2 repositories listed Syntology ran 1 of 7 samples · 6 unverified · 2 pointer-only (licence)However, in the paper, we reveal that CLIP is insensitive to different mask proposals and tends to produce similar predictions for various mask proposals of the same image.
-
20 Apr 2023 2 repositories listedWe conclude that SAM shows impressive zero-shot segmentation performance for certain medical imaging datasets, but moderate to poor performance for others.
-
12 Apr 2023 2 repositories listedThese phenomena conflict with conventional explainability methods based on the class attention map (CAM), where the raw model can highlight the local foreground regions using global supervision without alignment.
-
14 Mar 2023 2 repositories listedWe present OpenSeeD, a simple Open-vocabulary Segmentation and Detection framework that jointly learns from different segmentation and detection datasets.
-
16 Aug 2020 2 repositories listedIn this paper, we propose a novel context-aware feature generation method for zero-shot segmentation named CaGNet.
-
11 Jul 2025 1 repository listedDue to the excellent performance in yielding high-quality, zero-shot segmentation, Segment Anything Model (SAM) and its variants have been widely applied in diverse scenarios such as healthcare and intelligent…
-
3 Jun 2025 1 repository listedIn this paper, we investigate the efficacy of using a state-of-the-art image segmentation model, Segment Anything Model 2 (SAM2), in a zero-shot manner for individual tree detection and segmentation.
-
13 May 2025 1 repository listedAs AI-generated imagery becomes ubiquitous, invisible watermarks have emerged as a primary line of defense for copyright and provenance.
-
16 Apr 2025 1 repository listedExisting zero-shot 3D point cloud segmentation methods often struggle with limited transferability from seen classes to unseen classes and from semantic to visual space.
-
9 Jan 2025 1 repository listedDigital Pathology is a cornerstone in the diagnosis and treatment of diseases.
-
6 Dec 2024 1 repository listedIn the context of medical Augmented Reality (AR) applications, object tracking is a key challenge and requires a significant amount of annotation masks.
-
5 Dec 2024 1 repository listedImage segmentation foundation models (SFMs) like Segment Anything Model (SAM) have achieved impressive zero-shot and interactive segmentation across diverse domains.
-
2 Dec 2024 1 repository listedVision-Language Models (VLMs) trained with contrastive loss have achieved significant advancements in various vision and language tasks.
-
6 Nov 2024 1 repository listedWe present 3DGS-CD, the first 3D Gaussian Splatting (3DGS)-based method for detecting physical object rearrangements in 3D scenes.
-
1 Nov 2024 1 repository listedTo address this limitation, we propose a novel zero-shot image matting model, called ZIM, with two key contributions: First, we develop a label converter that transforms segmentation labels into detailed matte labels,…
-
3 Oct 2024 1 repository listed Syntology ran 2 of 3 samples · 1 unverifiedWe investigate the internal representations of vision-language models (VLMs) to address hallucinations, a persistent challenge despite advances in model size and training.
-
4 Sep 2024 1 repository listedSegment Anything Model (SAM) has demonstrated powerful zero-shot segmentation performance in natural scenes.
-
23 Aug 2024 1 repository listedThe unprecedented developments in segmentation foundational models have become a dominant force in the field of computer vision, introducing a multitude of previously unexplored capabilities in a wide range of natural…
-
19 Aug 2024 1 repository listedWe train SAM-UNet on SA-Med2D-16M, the largest 2-dimensional medical image segmentation dataset to date, yielding a universal pretrained model for medical images.
-
7 Aug 2024 1 repository listedThe framework consists of two main parts: a Single-Shot PCI Estimation Network and a Dense Captioning Network.
-
3 Aug 2024 1 repository listedThe Segment Anything Model 2 (SAM 2) is the latest generation foundation model for image and video segmentation.
-
22 Jul 2024 1 repository listedRapid and accurate diagnosis of pneumothorax, utilizing chest X-ray and computed tomography (CT), is crucial for assisted diagnosis.
-
18 Jul 2024 1 repository listedSpecifically, our model leverages the Segment Anything Model (SAM) model to segment the target regions from images rendered from the 3D shape.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections