Browse State-of-the-Art › Open-Vocabulary Semantic Segmentation
Open-Vocabulary Semantic Segmentation
54 papers with code · 0 benchmarks · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 0 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| ADE20K-150 (0 rows) | no rows in the archive | — | — | ||
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 54 papers with code (95 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
21 Mar 2023 3 repositories listed Syntology ran 5 of 6 samples · 1 unverifiedOpen-vocabulary semantic segmentation presents the challenge of labeling each pixel within an image based on a wide range of text descriptions.
-
23 Feb 2023 3 repositories listed Syntology ran 1 of 9 samples · 8 unverifiedA side network is attached to a frozen CLIP model with two branches: one for predicting mask proposals, and the other for predicting attention bias which is applied in the CLIP model to recognize the class of masks.
-
2 Oct 2024 2 repositories listed Syntology ran 3 of 9 samples · 6 unverified · 9 pointer-only (licence)To tackle this issue, we propose a simple and general upsampler, SimFeatUp, to restore lost spatial information in deep features in a training-free style.
-
17 Jun 2024 2 repositories listed Syntology ran 11 of 15 samples · 4 unverifiedOpen-vocabulary part segmentation (OVPS) is an emerging research area focused on segmenting fine-grained entities using diverse and previously unseen vocabularies.
-
11 Sep 2023 2 repositories listedIn this paper, we propose to the best of our knowledge the first algorithm for open-vocabulary panoptic segmentation in 3D scenes.
-
12 Apr 2023 2 repositories listedThese phenomena conflict with conventional explainability methods based on the class attention map (CAM), where the raw model can highlight the local foreground regions using global supervision without alignment.
-
29 Dec 2021 2 repositories listedHowever, semantic segmentation and the CLIP model perform on different visual granularity, that semantic segmentation processes on pixels while CLIP performs on images.
-
26 Jun 2025 1 repository listedTraining-free open-vocabulary semantic segmentation (OVS) aims to segment images given a set of arbitrary textual categories without costly model fine-tuning.
-
11 Jun 2025 1 repository listedOpen-Vocabulary semantic segmentation (OVSS) and domain generalization in semantic segmentation (DGSS) highlight a subtle complementarity that motivates Open-Vocabulary Domain-Generalized Semantic Segmentation (OV-DGSS).
-
28 May 2025 1 repository listedRecently, test-time adaptation has attracted wide interest in the context of vision-language models for image classification.
-
22 May 2025 1 repository listedTo the best of our knowledge, OpenSeg-R is the first framework to introduce explicit step-by-step visual reasoning into OVS.
-
14 Apr 2025 1 repository listedOur plug-and-play method, coined FLOSS, is orthogonal and complementary to existing OVSS methods, offering a ''free lunch'' to systematically improve OVSS without labels and additional training.
-
27 Mar 2025 1 repository listedOpen-vocabulary semantic segmentation models associate vision and text to label pixels from an undefined set of classes using textual queries, providing versatile performance on novel datasets.
-
25 Mar 2025 1 repository listed Syntology ran 0 of 5 samples · 5 unverifiedWe propose a training-free method for open-vocabulary semantic segmentation using Vision-and-Language Models (VLMs).
-
29 Jan 2025 1 repository listedOpen-vocabulary semantic segmentation (OVSS) is an open-world task that aims to assign each pixel within an image to a specific class defined by arbitrary text descriptions.
-
16 Jan 2025 1 repository listed Syntology ran 7 of 11 samples · 4 unverifiedOpen-Vocabulary Part Segmentation (OVPS) is an emerging field for recognizing fine-grained parts in unseen categories.
-
1 Jan 2025 1 repository listedThe core of FGAseg is a Pixel-Level Alignment module that employs a cross-modal attention mechanism and a text-pixel alignment loss to refine the coarse-grained alignment from CLIP, achieving finer-grained pixel-text…
-
20 Dec 2024 1 repository listedSelf-supervised visual foundation models produce powerful embeddings that achieve remarkable performance on a wide range of downstream tasks.
-
16 Dec 2024 1 repository listedOpen-vocabulary image segmentation has been advanced through the synergy between mask generators and vision-language models like Contrastive Language-Image Pre-training (CLIP).
-
5 Dec 2024 1 repository listedMask-Adapter integrates seamlessly into open-vocabulary segmentation methods based on mask pooling in a plug-and-play manner, delivering more accurate classification results.
-
21 Nov 2024 1 repository listedThe proposed CLIPer includes an early-layer fusion module and a fine-grained compensation module.
-
20 Nov 2024 1 repository listed Syntology ran 0 of 11 samples · 11 unverifiedIn our approach, we developed a mask generator based on the denoising UNet from a pre-trained diffusion model, leveraging its capability for precise textual control over dense pixel representations and enhancing the…
-
18 Nov 2024 1 repository listedRecent advances in foundational Vision Language Models (VLMs) have reshaped the evaluation paradigm in computer vision tasks.
-
15 Nov 2024 1 repository listed Syntology ran 0 of 9 samples · 9 unverified · 9 pointer-only (licence)Open-vocabulary semantic segmentation aims to assign semantic labels to each pixel without relying on a predefined set of categories.
-
18 Aug 2024 1 repository listedIn this paper, we introduce OVOSE, the first Open-Vocabulary Semantic Segmentation algorithm for Event cameras.
-
9 Aug 2024 1 repository listedWe present lazy visual grounding, a two-stage approach of unsupervised object mask discovery followed by object grounding, for open-vocabulary semantic segmentation.
-
9 Aug 2024 1 repository listed Syntology ran 3 of 9 samples · 6 unverified · 9 pointer-only (licence)ProxyCLIP leverages the spatial feature correspondence from VFMs as a form of proxy attention to augment CLIP, thereby inheriting the VFMs' robust local consistency and maintaining CLIP's exceptional zero-shot transfer…
-
11 Jul 2024 1 repository listed Syntology ran 3 of 5 samples · 2 unverifiedCLIP, as a vision-language model, has significantly advanced Open-Vocabulary Semantic Segmentation (OVSS) with its zero-shot capabilities.
-
3 Jul 2024 1 repository listed Syntology ran 7 of 7 samples · 0 unverifiedWe propose UniSeg3D, a unified 3D scene understanding framework that achieves panoptic, semantic, instance, interactive, referring, and open-vocabulary segmentation tasks within a single model.
-
14 Jun 2024 1 repository listed Syntology ran 3 of 12 samples · 9 unverifiedTo learn a consistent semantic structure from CLIP, the SSC Loss aligns the inter-classes affinity in the image feature space with that in the text feature space of CLIP, thereby improving the generalization ability of…
Syntology lines on 12 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections