Browse State-of-the-Art › Unsupervised Semantic Segmentation with Language-image Pre-training
Unsupervised Semantic Segmentation with Language-image Pre-training
14 papers with code · 12 benchmarks · 7 datasets archive 2025-07-28
A segmentation task which does not utilise any human-level supervision for semantic segmentation except for a backbone which is initialised with features pre-trained with image-level labels.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
12 leaderboard tables shown for this task, 12 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 12 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
7 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
14 shown of 14 papers with code (14 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
22 Feb 2022 6 repositories listedWith only text supervision and without any pixel-level annotations, GroupViT learns to group together semantic regions and successfully transfers to the task of semantic segmentation in a zero-shot manner, i.
-
18 Oct 2022 2 repositories listedIn this work we examine how well vision-language models are able to understand where objects reside within an image and group together visually related parts of the imagery.
-
14 Jun 2022 2 repositories listedSemantic segmentation has a broad range of applications, but its real-world impact has been significantly limited by the prohibitive annotation costs necessary to enable deployment.
-
29 May 2025 1 repository listedImage-text models excel at image-level tasks but struggle with detailed visual understanding.
-
2 Dec 2024 1 repository listedVision-Language Models (VLMs) trained with contrastive loss have achieved significant advancements in various vision and language tasks.
-
18 Nov 2024 1 repository listedRecent advances in foundational Vision Language Models (VLMs) have reshaped the evaluation paradigm in computer vision tasks.
-
15 Nov 2024 1 repository listed Syntology ran 0 of 9 samples · 9 unverified · 9 pointer-only (licence)Open-vocabulary semantic segmentation aims to assign semantic labels to each pixel without relying on a predefined set of categories.
-
Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary Segmentation14 Nov 2024 1 repository listedSpecifically, we introduce Trident, a training-free framework that first splices features extracted by CLIP and DINO from sub-images, then leverages SAM's encoder to create a correlation matrix for global aggregation,…
-
9 Aug 2024 1 repository listed Syntology ran 3 of 9 samples · 6 unverified · 9 pointer-only (licence)ProxyCLIP leverages the spatial feature correspondence from VFMs as a form of proxy attention to augment CLIP, thereby inheriting the VFMs' robust local consistency and maintaining CLIP's exceptional zero-shot transfer…
-
30 Mar 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)We identify a critical bias in contemporary CLIP-based models, which we denote as single tag bias.
-
21 Dec 2023 1 repository listedThe crux of learning vision-language models is to extract semantically aligned information from visual and linguistic data.
-
20 Dec 2023 1 repository listedAs a result, we dissect the preservation of patch-wise spatial information in CLIP and proposed a local-to-global framework to obtain image tags.
-
1 Dec 2022 1 repository listed Syntology ran 1 of 8 samples · 7 unverifiedExisting open-world segmentation methods have shown impressive advances by employing contrastive learning (CL) to learn diverse visual concepts and transferring the learned image-level understanding to the segmentation…
-
2 Dec 2021 1 repository listedContrastive Language-Image Pre-training (CLIP) has made a remarkable breakthrough in open-vocabulary zero-shot image recognition.
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections