Datasets › Cityscapes

Cityscapes

Introduced by Marius Cordts et al. in The Cityscapes Dataset for Semantic Urban Scene Understanding1 Jan 2016 archive 2025-07-28

Cityscapes is a large-scale database which focuses on semantic understanding of urban street scenes. It provides semantic, instance-wise, and dense pixel annotations for 30 classes grouped into 8 categories (flat surfaces, humans, vehicles, constructions, objects, nature, sky, and void). The dataset consists of around 5000 fine annotated images and 20000 coarse annotated ones. Data was captured in 50 cities during several months, daytimes, and good weather conditions. It was originally recorded as video so the frames were manually selected to have the following features: large number of dynamic objects, varying scene layout, and varying background.

Source: A Review on Deep Learning Techniques Applied to Semantic Segmentation Image Source: https://www.cityscapes-dataset.com/dataset-overview/

Benchmarks archive 2025-07-28

All 51 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Semantic Segmentation Cityscapes test VLTSeg Mean IoU (class) 86.4 Strong but simple: A Baseline for Domain Generalized... VLTSeg/VLTSeg 105 Compare
Semantic Segmentation Cityscapes val ViT-P (InternImage-H) mIoU 87.4 The Missing Point in Vision Transformers for Universal... sajjad-sh33/vit-p 99 Compare
Real-Time Semantic Segmentation Cityscapes test PIDNet-L mIoU 80.6% PIDNet: A Real-time Semantic Segmentation Network... XuJiacong/PIDNet +5 39 Compare
Panoptic Segmentation Cityscapes val ViT-P (OneFormer, InternImage-H) PQ 70.8 The Missing Point in Vision Transformers for Universal... sajjad-sh33/vit-p 37 Compare
Semi-Supervised Semantic Segmentation Cityscapes 12.5% labeled UniMatch V2 (DINOv2-B) Validation mIoU 84.3% UniMatch V2: Pushing the Limit of Semi-Supervised... LiheYoung/UniMatch-V2 33 Compare
Semi-Supervised Semantic Segmentation Cityscapes 25% labeled UniMatch V2 (DINOv2-B) Validation mIoU 84.5% UniMatch V2: Pushing the Limit of Semi-Supervised... LiheYoung/UniMatch-V2 30 Compare
Real-Time Semantic Segmentation Cityscapes val PIDNet-L mIoU 80.9% PIDNet: A Real-time Semantic Segmentation Network... XuJiacong/PIDNet +5 24 Compare
Semi-Supervised Semantic Segmentation Cityscapes 50% labeled UniMatch V2 (DINOv2-B) Validation mIoU 85.1% UniMatch V2: Pushing the Limit of Semi-Supervised... LiheYoung/UniMatch-V2 23 Compare
Image-to-Image Translation Cityscapes Labels-to-Photo DP-SIMS (ConvNext-L) mIoU 76.3 Unlocking Pre-trained Image Backbones for Semantic Image... — 21 Compare
Semi-Supervised Semantic Segmentation Cityscapes 6.25% labeled UniMatch V2 (DINOv2-B) Validation mIoU 83.6 UniMatch V2: Pushing the Limit of Semi-Supervised... LiheYoung/UniMatch-V2 18 Compare
Instance Segmentation Cityscapes val ViT-P (OneFormer, ConvNeXt-L, single-scale, 512x1024, Mapillary Vistas-pretrained) mask AP 49.0 The Missing Point in Vision Transformers for Universal... sajjad-sh33/vit-p 17 Compare
Unsupervised Semantic Segmentation Cityscapes test CUPS mIoU 26.8 Scene-Centric Unsupervised Panoptic Segmentation visinf/cups 14 Compare
Robust Object Detection Cityscapes FGT (SD-1.5 Backbone) mPC [AP] 27.4 Boosting Domain Generalized and Adaptive Detection with... heboyong/fitness-generalization-transferability 13 Compare
Semi-Supervised Semantic Segmentation Cityscapes 100 samples labeled SemiVL (ViT-B/16) Validation mIoU 76.2 SemiVL: Semi-Supervised Semantic Segmentation with... google-research/semivl 13 Compare
Unsupervised Semantic Segmentation with Language-image Pre-training Cityscapes val CorrCLIP mIoU 51.1 CorrCLIP: Reconstructing Correlations in CLIP with... zdk258/CorrCLIP 12 Compare
Instance Segmentation Cityscapes test Deep Watershed Transform Deep Watershed Transform for Instance Segmentation min2209/dwt +1 11 Compare
Panoptic Segmentation Cityscapes test OneFormer (ConvNeXt-L, single-scale, Mapillary Vistas-Pretrained) PQ 68.0 OneFormer: One Transformer to Rule Universal Image Segmentation huggingface/transformers +3 10 Compare
Federated Learning Cityscapes heterogeneous SiloBN + ASAM mIoU 49.75 Improving Generalization in Federated Learning by... debcaldarola/fedsam 9 Compare
Video Semantic Segmentation Cityscapes val TMANet-50 mIoU 80.3 Temporal Memory Attention for Video Semantic Segmentation wanghao9610/TMANet 9 Compare
Image Generation Cityscapes Projected GAN FID-10k-training-steps 3.41 Projected GANs Converge Faster autonomousvision/projected_gan +2 6 Compare
Image-to-Image Translation Cityscapes Photo-to-Labels pix2pix Class IOU 0.32 Image-to-Image Translation with Conditional Adversarial Networks tensorflow/models +191 5 Compare
Open Vocabulary Semantic Segmentation Cityscapes FC-CLIP mIoU 56.2 Convolutions Die Hard: Open-Vocabulary Segmentation with... bytedance/fc-clip 5 Compare
Unsupervised Panoptic Segmentation Cityscapes CUPS (54 pseudo-classes) PQ 30.6 Scene-Centric Unsupervised Panoptic Segmentation visinf/cups 5 Compare
Video Prediction Cityscapes 128x128 GHVAEs FVD 418.00 ± 5.0 Greedy Hierarchical Variational Autoencoders for... — 5 Compare
Monocular Depth Estimation Cityscapes SwinMTL RMSE 5.481 SwinMTL: A Shared Architecture for Simultaneous Depth... pardistaghavi/swinmtl 3 Compare
Multi-Task Learning Cityscapes test SwinMTL mIoU 76.41 SwinMTL: A Shared Architecture for Simultaneous Depth... pardistaghavi/swinmtl 3 Compare
Semi-Supervised Semantic Segmentation Cityscapes 2% labeled GIST and RIST (DeepLabv2 with ResNet101, MSCOCO pre-trained) Validation mIoU 53.51% The GIST and RIST of Iterative Self-Training for... — 3 Compare
Semi-Supervised Semantic Segmentation Cityscapes 5% labeled GIST and RIST (DeepLabv2 with ResNet101, MSCOCO pre-trained) Validation mIoU 59.98% The GIST and RIST of Iterative Self-Training for... — 3 Compare
Semi-Supervised Semantic Segmentation Cityscapes 93 labeled AEL (DeepLab v3+ with ResNet-101 pretraind on ImageNet-1K) Validation mIoU 74.28 Semi-Supervised Semantic Segmentation via Adaptive... hzhupku/semiseg-ael 3 Compare
Video Prediction Cityscapes DMVFN LPIPS 0.0558 A Dynamic Multi-Scale Voxel Flow Network for Video Prediction megvii-research/CVPR2023-DMVFN 3 Compare
Depth Estimation Cityscapes test SwinMTL RMSE 6.352 SwinMTL: A Shared Architecture for Simultaneous Depth... pardistaghavi/swinmtl 2 Compare
Edge Detection Cityscapes test RPCNet AP 86.15% Joint Semantic Segmentation and Boundary Detection using... — 2 Compare
Real-Time Semantic Segmentation Cityscapes S^2-FPN34 mIoU 77.4 S²-FPN: Scale-ware Strip Attention Guided Feature... mohamedac29/s2-fpn 2 Compare
Robust Object Detection Cityscapes test Faster R-CNN with Stylized Training Data mPC [AP] 17.2 Benchmarking Robustness in Object Detection: Autonomous... bethgelab/imagecorruptions +3 2 Compare
Semantic Segmentation Cityscapes SPFNet34M mIoU 77.8 S²-FPN: Scale-ware Strip Attention Guided Feature... mohamedac29/s2-fpn 2 Compare
Semi-Supervised Semantic Segmentation Cityscapes with extra (no coarse labels) Dense FixMatch (DeepLabv3+ ResNet-101, over-sampling, single pass eval) Validation mIoU 80.82 Dense FixMatch: a simple semi-supervised learning method... miquelmarti/DenseFixMatch 2 Compare
2D Semantic Segmentation Cityscapes val SERNet-Former mIoU 87.35 SERNet-Former: Semantic Segmentation by Efficient... serdarch/sernet-former +1 1 Compare
Image Generation Cityscapes-25K 256x512 SB-GAN FID 62.97 Semantic Bottleneck Scene Generation azadis/SB-GAN +1 1 Compare
Image Generation Cityscapes-5K 256x512 SB-GAN FID 65.49 Semantic Bottleneck Scene Generation azadis/SB-GAN +1 1 Compare
Instance Segmentation Cityscapes CAST AP 33.9 CAST: Contrastive Adaptation and Distillation for... — 1 Compare
Interactive Segmentation Cityscapes val IOG Instance Average IoU 83.8 Interactive Object Segmentation With Inside-Outside Guidance shiyinzhang/Inside-Outside-Guidance +1 1 Compare
Knowledge Distillation Cityscapes CAST AP 33.9 CAST: Contrastive Adaptation and Distillation for... — 1 Compare
Overlapped 10-1 Cityscapes MiB+AWT mIoU 44.9 Attribution-aware Weight Transfer: A Warm-Start... dfki-av/awt-for-ciss 1 Compare
Overlapped 14-1 Cityscapes MiB+AWT mIoU 46.9 Attribution-aware Weight Transfer: A Warm-Start... dfki-av/awt-for-ciss 1 Compare
Real-time Instance Segmentation Cityscapes test CenterPoly AP 15.54 CenterPoly: real-time instance segmentation using... hu64/centerpoly 1 Compare
Scene Parsing Cityscapes test VCD No Coarse mIoU 82.3 Variational Context-Deformable ConvNets for Indoor Scene Parsing — 1 Compare
Semi-Supervised Instance Segmentation Cityscapes CAST AP 33.9 CAST: Contrastive Adaptation and Distillation for... — 1 Compare
Semi-Supervised Semantic Segmentation Cityscapes 10% labeled IM++ (416x208, 2.7m parameters, no pretraining) Mean IoU (class) 0.428 Inconsistency Masks: Removing the Uncertainty from... michaelvorndran/inconsistencymasks 1 Compare
Unsupervised Semantic Segmentation Cityscapes val Segmenter ViT-S/16 mIoU 21.8 Drive&Segment: Unsupervised Semantic Segmentation of... vobecant/DriveAndSegment 1 Compare
Weakly-Supervised Semantic Segmentation Cityscapes val CARB mIoU 52.1 Weakly Supervised Semantic Segmentation for Driving Scenes k0u-id/carb 1 Compare
Weakly-Supervised Semantic Segmentation Cityscapes test CARB mIoU 51.8 Weakly Supervised Semantic Segmentation for Driving Scenes k0u-id/carb 1 Compare

Papers archive 2025-07-28

30 shown of 298 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 3,702. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Boosting Domain Generalized and Adaptive Detection with Diffusion Models: Fitness, Generalization, and Transferability 1 2 26 Jun 2025 not harvested
FARCLUSS: Fuzzy Adaptive Rebalancing and Contrastive Uncertainty Learning for Semi-Supervised Semantic Segmentation 1 4 11 Jun 2025 not harvested
CAST: Contrastive Adaptation and Distillation for Semi-Supervised Instance Segmentation 0 3 28 May 2025 not harvested
The Missing Point in Vision Transformers for Universal Image Segmentation 1 3 26 May 2025 not harvested
Scene-Centric Unsupervised Panoptic Segmentation 1 4 2 Apr 2025 ran 0 of 8 samples (8 unverified)
Your ViT is Secretly an Image Segmentation Model 1 1 24 Mar 2025 ran 1 of 2 samples (1 unverified)
Generalized Diffusion Detector: Mining Robust Features from Diffusion Models for Domain-Generalized Detection 1 2 3 Mar 2025 not harvested
VPNeXt -- Rethinking Dense Decoding for Plain Vision Transformer 0 1 23 Feb 2025 not harvested
Confidence-Weighted Boundary-Aware Learning for Semi-Supervised Semantic Segmentation 1 4 21 Feb 2025 not harvested
PhysAug: A Physical-guided and Frequency-based Data Augmentation for Single-Domain Generalized Object Detection 1 1 16 Dec 2024 not harvested
GraPix: Exploring Graph Modularity Optimization for Unsupervised Pixel Clustering 1 3 4 Dec 2024 not harvested
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training 1 1 2 Dec 2024 not harvested
CorrCLIP: Reconstructing Correlations in CLIP with Off-the-Shelf Foundation Models for Open-Vocabulary Semantic Segmentation 1 1 15 Nov 2024 ran 0 of 9 samples (9 unverified; 9 pointer-only for licence)
Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary Segmentation 1 1 14 Nov 2024 not harvested
Rethinking Decoders for Transformer-based Semantic Segmentation: A Compression Perspective 1 2 5 Nov 2024 not harvested
UniMatch V2: Pushing the Limit of Semi-Supervised Semantic Segmentation 1 4 14 Oct 2024 ran 2 of 5 samples (3 unverified)
ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation 1 1 9 Aug 2024 ran 3 of 9 samples (6 unverified; 9 pointer-only for licence)
CSFNet: A Cosine Similarity Fusion Network for Real-Time RGB-X Semantic Segmentation of Driving Scenes 1 4 1 Jul 2024 not harvested
DSNet: A Novel Way to Use Atrous Convolutions in Semantic Segmentation 1 3 6 Jun 2024 not harvested
Revisiting and Maximizing Temporal Knowledge in Semi-supervised Semantic Segmentation 1 8 31 May 2024 ran 6 of 11 samples (5 unverified)
Boosting Unsupervised Semantic Segmentation with Principal Mask Proposals 1 2 25 Apr 2024 not harvested
TTD: Text-Tag Self-Distillation Enhancing Image-Text Alignment in CLIP to Alleviate Single Tag Bias 1 4 30 Mar 2024 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
SwinMTL: A Shared Architecture for Simultaneous Depth Estimation and Semantic Segmentation from Monocular Camera Images 1 6 15 Mar 2024 not harvested
EAGLE: Eigen Aggregation Learning for Object-Centric Unsupervised Semantic Segmentation 1 2 3 Mar 2024 ran 14 of 19 samples (5 unverified)
SERNet-Former: Semantic Segmentation by Efficient Residual Network with Attention-Boosting Gates and Attention-Fusion Networks 2 3 28 Jan 2024 not harvested
Inconsistency Masks: Removing the Uncertainty from Input-Pseudo-Label Pairs 1 1 25 Jan 2024 not harvested
Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data 7 2 19 Jan 2024 ran 3 of 11 samples (8 unverified)
Fully Attentional Networks with Self-emerging Token Labeling 1 1 8 Jan 2024 not harvested
Unsupervised Universal Image Segmentation 2 1 28 Dec 2023 ran 13 of 15 samples (2 unverified)
Manydepth2: Motion-Aware Self-Supervised Multi-Frame Monocular Depth Estimation in Dynamic Scenes 1 1 23 Dec 2023 ran 0 of 9 samples (9 unverified)

The full list of 298 is in the JSON twin.

Dataset loaders archive 2025-07-28

8 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • Semi-Supervised Semantic Segmentation on Cityscapes 6.25% labeled
  • Semi-Supervised Semantic Segmentation on Cityscapes 12.5% labeled
  • Cityscapes with extra (no coarse labels)
  • Cityscapes with extra (no coarse)
  • Cityscapes heterogeneous
  • Cityscapes 6.25% labeled
  • Cityscapes 5% labeled
  • Cityscapes 2% labeled
  • Cityscapes 128x128
  • Cityscapes 93 labeled
  • Cityscapes 10% labeled
  • Cityscapes
  • Cityscapes-5K 256x512
  • Cityscapes-25K 256x512
  • Cityscapes val
  • Cityscapes test
  • Cityscapes Photo-to-Labels
  • Cityscapes Labels-to-Photo
  • Cityscapes 50% labeled
  • Cityscapes 25% labeled
  • Cityscapes 12.5% labeled
  • Cityscapes 100 samples labeled

22 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections