| Semantic Segmentation |
ADE20K |
ViT-P (InternImage-H) Validation mIoU 63.6 |
The Missing Point in Vision Transformers for Universal... |
sajjad-sh33/vit-p |
235 |
Compare |
| Semantic Segmentation |
ADE20K val |
BEiT-3 mIoU 62.8 |
Image as a Foreign Language: BEiT Pretraining for All... |
microsoft/unilm +1 |
95 |
Compare |
| Panoptic Segmentation |
ADE20K val |
OneFormer (InternImage-H, emb_dim=256, single-scale, 896x896) PQ 54.5 |
OneFormer: One Transformer to Rule Universal Image Segmentation |
huggingface/transformers +3 |
25 |
Compare |
| Open Vocabulary Semantic Segmentation |
ADE20K-150 |
Mask-Adapter mIoU 38.2 |
Mask-Adapter: The Devil is in the Masks for... |
hustvl/maskadapter |
23 |
Compare |
| Open Vocabulary Semantic Segmentation |
ADE20K-847 |
UMG-CLIP-E/14 mIoU 17.3 |
UMG-CLIP: A Unified Multi-Granularity Vision Generalist... |
lygsbw/umg-clip |
19 |
Compare |
| Image-to-Image Translation |
ADE20K Labels-to-Photos |
DP-SIMS (ConvNext-L) mIoU 54.3 |
Unlocking Pre-trained Image Backbones for Semantic Image... |
— |
16 |
Compare |
| Instance Segmentation |
ADE20K val |
OneFormer (InternImage-H, emb_dim=1024, single-scale, 896x896, COCO-Pretrained) AP 44.2 |
OneFormer: One Transformer to Rule Universal Image Segmentation |
huggingface/transformers +3 |
14 |
Compare |
| Unsupervised Semantic Segmentation with Language-image Pre-training |
ADE20K |
CorrCLIP Mean IoU (val) 30.7 |
CorrCLIP: Reconstructing Correlations in CLIP with... |
zdk258/CorrCLIP |
13 |
Compare |
| Open Vocabulary Panoptic Segmentation |
ADE20K |
UMG-CLIP-E/14 PQ 31.6 |
UMG-CLIP: A Unified Multi-Granularity Vision Generalist... |
lygsbw/umg-clip |
10 |
Compare |
| Overlapped 100-5 |
ADE20K |
MBS mIoU 42.8 |
Mitigating Background Shift in Class-Incremental... |
roadonep/eccv2024_mbs |
8 |
Compare |
| Image-to-Image Translation |
ADE20K-Outdoor Labels-to-Photos |
DP-GAN mIoU 40.4 |
Dual Pyramid Generative Adversarial Networks for... |
sj-li/dp_gan |
7 |
Compare |
| Overlapped 100-50 |
ADE20K |
MBS mIoU 45.7 |
Mitigating Background Shift in Class-Incremental... |
roadonep/eccv2024_mbs |
7 |
Compare |
| Overlapped 50-50 |
ADE20K |
MBS mIoU 45.4 |
Mitigating Background Shift in Class-Incremental... |
roadonep/eccv2024_mbs |
7 |
Compare |
| Overlapped 100-10 |
ADE20K |
MBS Mean IoU (test) 44.5 |
Mitigating Background Shift in Class-Incremental... |
roadonep/eccv2024_mbs |
6 |
Compare |
| Semi-Supervised Semantic Segmentation |
ADE20K 1/32 labeled |
UniMatch V2 Validation mIoU 45.0 |
UniMatch V2: Pushing the Limit of Semi-Supervised... |
LiheYoung/UniMatch-V2 |
5 |
Compare |
| Semi-Supervised Semantic Segmentation |
ADE20K 1/16 labeled |
UniMatch V2 Validation mIoU 46.7 |
UniMatch V2: Pushing the Limit of Semi-Supervised... |
LiheYoung/UniMatch-V2 |
5 |
Compare |
| Sound Prompted Semantic Segmentation |
ADE20K |
DenseAV mAP 32.7 |
Separating the "Chirp" from the "Chat": Self-supervised... |
mhamilton723/DenseAV |
4 |
Compare |
| Speech Prompted Semantic Segmentation |
ADE20K |
DenseAV mAP 48.7 |
Separating the "Chirp" from the "Chat": Self-supervised... |
mhamilton723/DenseAV |
4 |
Compare |
| Continual Semantic Segmentation |
ADE20K |
LGKD mIoU 37.5 |
Label-Guided Knowledge Distillation for Continual... |
ze-yang/lgkd |
2 |
Compare |
| Face Detection |
ADE20K |
CASSOD mIoU 42.86 |
CASSOD-Net: Cascaded and Separable Structures of Dilated... |
— |
1 |
Compare |
| Overlapped 25-25 |
ADE20K |
SATS-M Mean IoU (test) 32.56 |
SATS: Self-Attention Transfer for Continual Semantic Segmentation |
QIU023/SATS_Continual_Semantic_Seg |
1 |
Compare |
| Panoptic Segmentation |
ADE20K |
MasQCLIP PQ 23.3 |
MasQCLIP for Open-Vocabulary Universal Image Segmentation |
mlpc-ucsd/MasQCLIP |
1 |
Compare |
| Pose Transfer |
ADE20K |
SCAM FID 27.5 |
SCAM! Transferring humans between images with Semantic... |
nicolas-dufour/SCAM |
1 |
Compare |
| Reconstruction |
ADE20K |
SCAM PSNR 20 |
SCAM! Transferring humans between images with Semantic... |
nicolas-dufour/SCAM |
1 |
Compare |
| Scene Recognition |
ADE20K |
Semantic-Aware Scene Recogniton (ResNet-18) Top 1 Accuracy 62.55 |
Semantic-Aware Scene Recognition |
vpulab/Semantic-Aware-Scene-Recognition |
1 |
Compare |
| Scene Understanding |
ADE20K val |
CPN(ResNet-101) Mean IoU 46.3 |
Context Prior for Scene Segmentation |
ycszen/ContextPrior +1 |
1 |
Compare |
| Semi-Supervised Instance Segmentation |
ADE20K |
CAST AP 16.7 |
CAST: Contrastive Adaptation and Distillation for... |
— |
1 |
Compare |
| Weakly-Supervised Semantic Segmentation |
ADE20K val |
DHR (Swin-L, Mask2Former) mIoU 32.9 |
DHR: Dual Features-Driven Hierarchical Rebalancing in... |
shjo-april/DHR |
1 |
Compare |
| Zero-Shot Semantic Segmentation |
ADE20K-847 |
MAFT unseen mIoU 8.7 |
— |
— |
1 |
Compare |
| Open-Vocabulary Semantic Segmentation |
ADE20K-150 |
no rows |
— |
— |
0 |
Compare |
| Open Vocabulary Semantic Segmentation |
ADE20K-150 |
no rows |
— |
— |
0 |
Compare |
| Semantic Segmentation |
ADE20K-150 |
no rows |
— |
— |
0 |
Compare |