Papers › UMG-CLIP: A Unified Multi-Granularity Vision Generalist for Open-World Understanding

UMG-CLIP: A Unified Multi-Granularity Vision Generalist for Open-World Understanding

12 Jan 2024arXiv:2401.06397archive 2025-07-28

Bowen Shi, Peisen Zhao, Zichen Wang, Yuhang Zhang, Yaoming Wang, Jin Li, Wenrui Dai, Junni Zou, Hongkai Xiong, Qi Tian, Xiaopeng Zhang

Vision-language foundation models, represented by Contrastive Language-Image Pre-training (CLIP), have gained increasing attention for jointly understanding both vision and textual tasks. However, existing approaches primarily focus on training models to match global image representations with textual descriptions, thereby overlooking the critical alignment between local regions and corresponding text tokens. This paper extends CLIP with multi-granularity alignment. Notably, we deliberately construct a new dataset comprising pseudo annotations at various levels of granularities, encompassing image-level, region-level as well as pixel-level captions and tags. Accordingly, we develop a Unified Multi-Granularity learning framework, termed UMG-CLIP, which simultaneously empowers the model with versatile perception abilities across different levels of detail. With parameter efficient tuning, UMG-CLIP surpasses current widely used CLIP variants and achieves state-of-the-art performance on diverse image understanding benchmarks, including open-world recognition, retrieval, semantic segmentation, and panoptic segmentation tasks. We believe that UMG-CLIP represents a valuable advancement in vision-language foundation models. The code is available at https://github.com/lygsbw/UMG-CLIP.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

lygsbw/umg-clip officialmentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Open Vocabulary Panoptic SegmentationOpen Vocabulary Semantic SegmentationPanoptic SegmentationRetrievalSegmentationSemantic Segmentation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Open Vocabulary Panoptic Segmentation ADE20K UMG-CLIP-E/14 PQ 31.6 #1 of 10 Archive leaderboard report
Open Vocabulary Panoptic Segmentation ADE20K UMG-CLIP-L/14 PQ 29.1 #3 of 10 Archive leaderboard report
Open Vocabulary Semantic Segmentation ADE20K-150 UMG-CLIP-E/14 mIoU 38.2 #3 of 23 Archive leaderboard report
Open Vocabulary Semantic Segmentation ADE20K-150 UMG-CLIP-L/14 mIoU 36.1 #7 of 23 Archive leaderboard report
Open Vocabulary Semantic Segmentation ADE20K-847 UMG-CLIP-E/14 mIoU 17.3 #1 of 19 Archive leaderboard report
Open Vocabulary Semantic Segmentation ADE20K-847 UMG-CLIP-L/14 mIoU 15.4 #5 of 19 Archive leaderboard report
Open Vocabulary Semantic Segmentation PASCAL Context-459 UMG-CLIP-E/14 mIoU 25.2 #2 of 15 Archive leaderboard report
Open Vocabulary Semantic Segmentation PASCAL Context-459 UMG-CLIP-L/14 mIoU 23.2 #5 of 15 Archive leaderboard report
Open Vocabulary Semantic Segmentation PASCAL Context-59 UMG-CLIP-L/14 mIoU 61.0 #6 of 24 Archive leaderboard report
Open Vocabulary Semantic Segmentation PascalVOC-20 UMG-CLIP-L/14 mIoU 97.9 #1 of 20 Archive leaderboard report
Open Vocabulary Semantic Segmentation PascalVOC-20b UMG-CLIP-E/14 mIoU 85.4 #1 of 4 Archive leaderboard report
Panoptic Segmentation COCO minival UMG-CLIP-E/14 AP 50.7 #4 of 31 Archive leaderboard report
Panoptic Segmentation COCO minival UMG-CLIP-E/14 PQ 59.5 #4 of 31 Archive leaderboard report
Panoptic Segmentation COCO minival UMG-CLIP-E/14 mIoU 69.7 #4 of 31 Archive leaderboard report
Panoptic Segmentation COCO minival UMG-CLIP-L/14 AP 49.7 #7 of 31 Archive leaderboard report
Panoptic Segmentation COCO minival UMG-CLIP-L/14 PQ 58.9 #7 of 31 Archive leaderboard report
Panoptic Segmentation COCO minival UMG-CLIP-L/14 mIoU 68.9 #7 of 31 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

CLIPFocus

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections