| Semantic Segmentation |
Cityscapes test |
VLTSeg Mean IoU (class) 86.4 |
Strong but simple: A Baseline for Domain Generalized... |
VLTSeg/VLTSeg |
105 |
Compare |
| Semantic Segmentation |
Cityscapes val |
ViT-P (InternImage-H) mIoU 87.4 |
The Missing Point in Vision Transformers for Universal... |
sajjad-sh33/vit-p |
99 |
Compare |
| Real-Time Semantic Segmentation |
Cityscapes test |
PIDNet-L mIoU 80.6% |
PIDNet: A Real-time Semantic Segmentation Network... |
XuJiacong/PIDNet +5 |
39 |
Compare |
| Panoptic Segmentation |
Cityscapes val |
ViT-P (OneFormer, InternImage-H) PQ 70.8 |
The Missing Point in Vision Transformers for Universal... |
sajjad-sh33/vit-p |
37 |
Compare |
| Semi-Supervised Semantic Segmentation |
Cityscapes 12.5% labeled |
UniMatch V2 (DINOv2-B) Validation mIoU 84.3% |
UniMatch V2: Pushing the Limit of Semi-Supervised... |
LiheYoung/UniMatch-V2 |
33 |
Compare |
| Semi-Supervised Semantic Segmentation |
Cityscapes 25% labeled |
UniMatch V2 (DINOv2-B) Validation mIoU 84.5% |
UniMatch V2: Pushing the Limit of Semi-Supervised... |
LiheYoung/UniMatch-V2 |
30 |
Compare |
| Real-Time Semantic Segmentation |
Cityscapes val |
PIDNet-L mIoU 80.9% |
PIDNet: A Real-time Semantic Segmentation Network... |
XuJiacong/PIDNet +5 |
24 |
Compare |
| Semi-Supervised Semantic Segmentation |
Cityscapes 50% labeled |
UniMatch V2 (DINOv2-B) Validation mIoU 85.1% |
UniMatch V2: Pushing the Limit of Semi-Supervised... |
LiheYoung/UniMatch-V2 |
23 |
Compare |
| Image-to-Image Translation |
Cityscapes Labels-to-Photo |
DP-SIMS (ConvNext-L) mIoU 76.3 |
Unlocking Pre-trained Image Backbones for Semantic Image... |
— |
21 |
Compare |
| Semi-Supervised Semantic Segmentation |
Cityscapes 6.25% labeled |
UniMatch V2 (DINOv2-B) Validation mIoU 83.6 |
UniMatch V2: Pushing the Limit of Semi-Supervised... |
LiheYoung/UniMatch-V2 |
18 |
Compare |
| Instance Segmentation |
Cityscapes val |
ViT-P (OneFormer, ConvNeXt-L, single-scale, 512x1024, Mapillary Vistas-pretrained) mask AP 49.0 |
The Missing Point in Vision Transformers for Universal... |
sajjad-sh33/vit-p |
17 |
Compare |
| Unsupervised Semantic Segmentation |
Cityscapes test |
CUPS mIoU 26.8 |
Scene-Centric Unsupervised Panoptic Segmentation |
visinf/cups |
14 |
Compare |
| Robust Object Detection |
Cityscapes |
FGT (SD-1.5 Backbone) mPC [AP] 27.4 |
Boosting Domain Generalized and Adaptive Detection with... |
heboyong/fitness-generalization-transferability |
13 |
Compare |
| Semi-Supervised Semantic Segmentation |
Cityscapes 100 samples labeled |
SemiVL (ViT-B/16) Validation mIoU 76.2 |
SemiVL: Semi-Supervised Semantic Segmentation with... |
google-research/semivl |
13 |
Compare |
| Unsupervised Semantic Segmentation with Language-image Pre-training |
Cityscapes val |
CorrCLIP mIoU 51.1 |
CorrCLIP: Reconstructing Correlations in CLIP with... |
zdk258/CorrCLIP |
12 |
Compare |
| Instance Segmentation |
Cityscapes test |
Deep Watershed Transform |
Deep Watershed Transform for Instance Segmentation |
min2209/dwt +1 |
11 |
Compare |
| Panoptic Segmentation |
Cityscapes test |
OneFormer (ConvNeXt-L, single-scale, Mapillary Vistas-Pretrained) PQ 68.0 |
OneFormer: One Transformer to Rule Universal Image Segmentation |
huggingface/transformers +3 |
10 |
Compare |
| Federated Learning |
Cityscapes heterogeneous |
SiloBN + ASAM mIoU 49.75 |
Improving Generalization in Federated Learning by... |
debcaldarola/fedsam |
9 |
Compare |
| Video Semantic Segmentation |
Cityscapes val |
TMANet-50 mIoU 80.3 |
Temporal Memory Attention for Video Semantic Segmentation |
wanghao9610/TMANet |
9 |
Compare |
| Image Generation |
Cityscapes |
Projected GAN FID-10k-training-steps 3.41 |
Projected GANs Converge Faster |
autonomousvision/projected_gan +2 |
6 |
Compare |
| Image-to-Image Translation |
Cityscapes Photo-to-Labels |
pix2pix Class IOU 0.32 |
Image-to-Image Translation with Conditional Adversarial Networks |
tensorflow/models +191 |
5 |
Compare |
| Open Vocabulary Semantic Segmentation |
Cityscapes |
FC-CLIP mIoU 56.2 |
Convolutions Die Hard: Open-Vocabulary Segmentation with... |
bytedance/fc-clip |
5 |
Compare |
| Unsupervised Panoptic Segmentation |
Cityscapes |
CUPS (54 pseudo-classes) PQ 30.6 |
Scene-Centric Unsupervised Panoptic Segmentation |
visinf/cups |
5 |
Compare |
| Video Prediction |
Cityscapes 128x128 |
GHVAEs FVD 418.00 ± 5.0 |
Greedy Hierarchical Variational Autoencoders for... |
— |
5 |
Compare |
| Monocular Depth Estimation |
Cityscapes |
SwinMTL RMSE 5.481 |
SwinMTL: A Shared Architecture for Simultaneous Depth... |
pardistaghavi/swinmtl |
3 |
Compare |
| Multi-Task Learning |
Cityscapes test |
SwinMTL mIoU 76.41 |
SwinMTL: A Shared Architecture for Simultaneous Depth... |
pardistaghavi/swinmtl |
3 |
Compare |
| Semi-Supervised Semantic Segmentation |
Cityscapes 2% labeled |
GIST and RIST (DeepLabv2 with ResNet101, MSCOCO pre-trained) Validation mIoU 53.51% |
The GIST and RIST of Iterative Self-Training for... |
— |
3 |
Compare |
| Semi-Supervised Semantic Segmentation |
Cityscapes 5% labeled |
GIST and RIST (DeepLabv2 with ResNet101, MSCOCO pre-trained) Validation mIoU 59.98% |
The GIST and RIST of Iterative Self-Training for... |
— |
3 |
Compare |
| Semi-Supervised Semantic Segmentation |
Cityscapes 93 labeled |
AEL (DeepLab v3+ with ResNet-101 pretraind on ImageNet-1K) Validation mIoU 74.28 |
Semi-Supervised Semantic Segmentation via Adaptive... |
hzhupku/semiseg-ael |
3 |
Compare |
| Video Prediction |
Cityscapes |
DMVFN LPIPS 0.0558 |
A Dynamic Multi-Scale Voxel Flow Network for Video Prediction |
megvii-research/CVPR2023-DMVFN |
3 |
Compare |
| Depth Estimation |
Cityscapes test |
SwinMTL RMSE 6.352 |
SwinMTL: A Shared Architecture for Simultaneous Depth... |
pardistaghavi/swinmtl |
2 |
Compare |
| Edge Detection |
Cityscapes test |
RPCNet AP 86.15% |
Joint Semantic Segmentation and Boundary Detection using... |
— |
2 |
Compare |
| Real-Time Semantic Segmentation |
Cityscapes |
S^2-FPN34 mIoU 77.4 |
S²-FPN: Scale-ware Strip Attention Guided Feature... |
mohamedac29/s2-fpn |
2 |
Compare |
| Robust Object Detection |
Cityscapes test |
Faster R-CNN with Stylized Training Data mPC [AP] 17.2 |
Benchmarking Robustness in Object Detection: Autonomous... |
bethgelab/imagecorruptions +3 |
2 |
Compare |
| Semantic Segmentation |
Cityscapes |
SPFNet34M mIoU 77.8 |
S²-FPN: Scale-ware Strip Attention Guided Feature... |
mohamedac29/s2-fpn |
2 |
Compare |
| Semi-Supervised Semantic Segmentation |
Cityscapes with extra (no coarse labels) |
Dense FixMatch (DeepLabv3+ ResNet-101, over-sampling, single pass eval) Validation mIoU 80.82 |
Dense FixMatch: a simple semi-supervised learning method... |
miquelmarti/DenseFixMatch |
2 |
Compare |
| 2D Semantic Segmentation |
Cityscapes val |
SERNet-Former mIoU 87.35 |
SERNet-Former: Semantic Segmentation by Efficient... |
serdarch/sernet-former +1 |
1 |
Compare |
| Image Generation |
Cityscapes-25K 256x512 |
SB-GAN FID 62.97 |
Semantic Bottleneck Scene Generation |
azadis/SB-GAN +1 |
1 |
Compare |
| Image Generation |
Cityscapes-5K 256x512 |
SB-GAN FID 65.49 |
Semantic Bottleneck Scene Generation |
azadis/SB-GAN +1 |
1 |
Compare |
| Instance Segmentation |
Cityscapes |
CAST AP 33.9 |
CAST: Contrastive Adaptation and Distillation for... |
— |
1 |
Compare |
| Interactive Segmentation |
Cityscapes val |
IOG Instance Average IoU 83.8 |
Interactive Object Segmentation With Inside-Outside Guidance |
shiyinzhang/Inside-Outside-Guidance +1 |
1 |
Compare |
| Knowledge Distillation |
Cityscapes |
CAST AP 33.9 |
CAST: Contrastive Adaptation and Distillation for... |
— |
1 |
Compare |
| Overlapped 10-1 |
Cityscapes |
MiB+AWT mIoU 44.9 |
Attribution-aware Weight Transfer: A Warm-Start... |
dfki-av/awt-for-ciss |
1 |
Compare |
| Overlapped 14-1 |
Cityscapes |
MiB+AWT mIoU 46.9 |
Attribution-aware Weight Transfer: A Warm-Start... |
dfki-av/awt-for-ciss |
1 |
Compare |
| Real-time Instance Segmentation |
Cityscapes test |
CenterPoly AP 15.54 |
CenterPoly: real-time instance segmentation using... |
hu64/centerpoly |
1 |
Compare |
| Scene Parsing |
Cityscapes test |
VCD No Coarse mIoU 82.3 |
Variational Context-Deformable ConvNets for Indoor Scene Parsing |
— |
1 |
Compare |
| Semi-Supervised Instance Segmentation |
Cityscapes |
CAST AP 33.9 |
CAST: Contrastive Adaptation and Distillation for... |
— |
1 |
Compare |
| Semi-Supervised Semantic Segmentation |
Cityscapes 10% labeled |
IM++ (416x208, 2.7m parameters, no pretraining) Mean IoU (class) 0.428 |
Inconsistency Masks: Removing the Uncertainty from... |
michaelvorndran/inconsistencymasks |
1 |
Compare |
| Unsupervised Semantic Segmentation |
Cityscapes val |
Segmenter ViT-S/16 mIoU 21.8 |
Drive&Segment: Unsupervised Semantic Segmentation of... |
vobecant/DriveAndSegment |
1 |
Compare |
| Weakly-Supervised Semantic Segmentation |
Cityscapes val |
CARB mIoU 52.1 |
Weakly Supervised Semantic Segmentation for Driving Scenes |
k0u-id/carb |
1 |
Compare |
| Weakly-Supervised Semantic Segmentation |
Cityscapes test |
CARB mIoU 51.8 |
Weakly Supervised Semantic Segmentation for Driving Scenes |
k0u-id/carb |
1 |
Compare |