Browse State-of-the-Art › Open Vocabulary Semantic Segmentation
Open Vocabulary Semantic Segmentation
71 papers with code · 14 benchmarks · 6 datasets archive 2025-07-28
Open-vocabulary semantic segmentation models aim to accurately assign a semantic label to each pixel in an image from a set of arbitrary open-vocabulary texts.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
15 leaderboard tables shown for this task, 14 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 15 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
6 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 71 papers with code (113 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
9 Mar 2025 3 repositories listedTraditional methods for reasoning segmentation rely on supervised fine-tuning with categorical labels and simple descriptions, limiting its out-of-domain generalization and lacking explicit reasoning processes.
-
21 Mar 2023 3 repositories listed Syntology ran 5 of 6 samples · 1 unverifiedOpen-vocabulary semantic segmentation presents the challenge of labeling each pixel within an image based on a wide range of text descriptions.
-
23 Feb 2023 3 repositories listed Syntology ran 1 of 9 samples · 8 unverifiedA side network is attached to a frozen CLIP model with two branches: one for predicting mask proposals, and the other for predicting attention bias which is applied in the CLIP model to recognize the class of masks.
-
2 Oct 2024 2 repositories listed Syntology ran 3 of 9 samples · 6 unverified · 9 pointer-only (licence)To tackle this issue, we propose a simple and general upsampler, SimFeatUp, to restore lost spatial information in deep features in a training-free style.
-
17 Jun 2024 2 repositories listed Syntology ran 11 of 15 samples · 4 unverifiedOpen-vocabulary part segmentation (OVPS) is an emerging research area focused on segmenting fine-grained entities using diverse and previously unseen vocabularies.
-
7 Dec 2023 2 repositories listed Syntology ran 5 of 10 samples · 5 unverified · 10 pointer-only (licence)We attribute this to the in-vocabulary embedding and domain-biased CLIP prediction.
-
30 Sep 2023 2 repositories listed Syntology ran 1 of 7 samples · 6 unverified · 2 pointer-only (licence)However, in the paper, we reveal that CLIP is insensitive to different mask proposals and tends to produce similar predictions for various mask proposals of the same image.
-
11 Sep 2023 2 repositories listedIn this paper, we propose to the best of our knowledge the first algorithm for open-vocabulary panoptic segmentation in 3D scenes.
-
12 Apr 2023 2 repositories listedThese phenomena conflict with conventional explainability methods based on the class attention map (CAM), where the raw model can highlight the local foreground regions using global supervision without alignment.
-
29 Dec 2021 2 repositories listedHowever, semantic segmentation and the CLIP model perform on different visual granularity, that semantic segmentation processes on pixels while CLIP performs on images.
-
26 Jun 2025 1 repository listedTraining-free open-vocabulary semantic segmentation (OVS) aims to segment images given a set of arbitrary textual categories without costly model fine-tuning.
-
11 Jun 2025 1 repository listedOpen-Vocabulary semantic segmentation (OVSS) and domain generalization in semantic segmentation (DGSS) highlight a subtle complementarity that motivates Open-Vocabulary Domain-Generalized Semantic Segmentation (OV-DGSS).
-
28 May 2025 1 repository listedRecently, test-time adaptation has attracted wide interest in the context of vision-language models for image classification.
-
22 May 2025 1 repository listedTo the best of our knowledge, OpenSeg-R is the first framework to introduce explicit step-by-step visual reasoning into OVS.
-
6 May 2025 1 repository listedPrompt engineering has shown remarkable success with large language models, yet its systematic exploration in computer vision remains limited.
-
14 Apr 2025 1 repository listedOur plug-and-play method, coined FLOSS, is orthogonal and complementary to existing OVSS methods, offering a ''free lunch'' to systematically improve OVSS without labels and additional training.
-
27 Mar 2025 1 repository listedOpen-vocabulary semantic segmentation models associate vision and text to label pixels from an undefined set of classes using textual queries, providing versatile performance on novel datasets.
-
25 Mar 2025 1 repository listed Syntology ran 0 of 5 samples · 5 unverifiedWe propose a training-free method for open-vocabulary semantic segmentation using Vision-and-Language Models (VLMs).
-
29 Jan 2025 1 repository listedOpen-vocabulary semantic segmentation (OVSS) is an open-world task that aims to assign each pixel within an image to a specific class defined by arbitrary text descriptions.
-
16 Jan 2025 1 repository listed Syntology ran 7 of 11 samples · 4 unverifiedOpen-Vocabulary Part Segmentation (OVPS) is an emerging field for recognizing fine-grained parts in unseen categories.
-
1 Jan 2025 1 repository listedThe core of FGAseg is a Pixel-Level Alignment module that employs a cross-modal attention mechanism and a text-pixel alignment loss to refine the coarse-grained alignment from CLIP, achieving finer-grained pixel-text…
-
20 Dec 2024 1 repository listedSelf-supervised visual foundation models produce powerful embeddings that achieve remarkable performance on a wide range of downstream tasks.
-
16 Dec 2024 1 repository listedOpen-vocabulary image segmentation has been advanced through the synergy between mask generators and vision-language models like Contrastive Language-Image Pre-training (CLIP).
-
5 Dec 2024 1 repository listedMask-Adapter integrates seamlessly into open-vocabulary segmentation methods based on mask pooling in a plug-and-play manner, delivering more accurate classification results.
-
26 Nov 2024 1 repository listed Syntology ran 7 of 17 samples · 10 unverifiedThis paper aims to address universal segmentation for image and video perception with the strong reasoning ability empowered by Visual Large Language Models (VLLMs).
-
21 Nov 2024 1 repository listedThe proposed CLIPer includes an early-layer fusion module and a fine-grained compensation module.
-
20 Nov 2024 1 repository listed Syntology ran 0 of 11 samples · 11 unverifiedIn our approach, we developed a mask generator based on the denoising UNet from a pre-trained diffusion model, leveraging its capability for precise textual control over dense pixel representations and enhancing the…
-
18 Nov 2024 1 repository listedRecent advances in foundational Vision Language Models (VLMs) have reshaped the evaluation paradigm in computer vision tasks.
-
15 Nov 2024 1 repository listed Syntology ran 0 of 9 samples · 9 unverified · 9 pointer-only (licence)Open-vocabulary semantic segmentation aims to assign semantic labels to each pixel without relying on a predefined set of categories.
-
18 Aug 2024 1 repository listedIn this paper, we introduce OVOSE, the first Open-Vocabulary Semantic Segmentation algorithm for Event cameras.
Syntology lines on 11 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections