Browse State-of-the-Art › Semantic Segmentation

Semantic Segmentation

6,644 papers with code · 150 benchmarks · 347 datasets archive 2025-07-28

Computer CodeComputer VisionMedicalRobots

Benchmarks archive 2025-07-28

151 leaderboard tables shown for this task (1 more in the archive withheld as spam; see /not-shown), 150 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 151 until expanded.

DatasetBest model (first row in archive order)PaperCodeSyntologyCompare
ADE20K (235 rows) ViT-P (InternImage-H) The Missing Point in Vision Transformers for Universal Image Segmentation code — Compare
NYU Depth v2 (121 rows) OmniVec2 OmniVec2 - A Novel Transformer based Network for Large Scale... — — Compare
Cityscapes test (105 rows) VLTSeg Strong but simple: A Baseline for Domain Generalized Dense... code — Compare
Cityscapes val (99 rows) ViT-P (InternImage-H) The Missing Point in Vision Transformers for Universal Image Segmentation code — Compare
ADE20K val (95 rows) BEiT-3 Image as a Foreign Language: BEiT Pretraining for All Vision and... code — Compare
PASCAL Context (66 rows) VPNeXt VPNeXt -- Rethinking Dense Decoding for Plain Vision Transformer — — Compare
S3DIS Area5 (61 rows) Sonata + PTv3 Sonata: Self-Supervised Learning of Reliable Point Representations code Syntology ran 5 of 19 samples · 14 unverified Compare
S3DIS (54 rows) Sonata + PTv3 Sonata: Self-Supervised Learning of Reliable Point Representations code Syntology ran 5 of 19 samples · 14 unverified Compare
PASCAL VOC 2012 test (51 rows) DeepLabv3+ (Xception-65-JFT) Encoder-Decoder with Atrous Separable Convolution for Semantic... code Syntology ran 43 of 72 samples · 29 unverified Compare
ScanNet (45 rows) DITR DINO in the Room: Leveraging 2D Foundation Models for 3D Segmentation code — Compare
SUN-RGBD (44 rows) GeminiFusion (Swin-Large) GeminiFusion: Efficient Pixel-wise Multimodal Fusion for Vision Transformer code Syntology ran 7 of 10 samples · 3 unverified Compare
DensePASS (36 rows) Trans4PASS+ (multi-scale) Behind Every Domain There is a Shift: Adapting Distortion-aware... code — Compare
PASCAL VOC 2012 val (29 rows) EfficientNet-L2+NAS-FPN (single scale test, with self-training) Rethinking Pre-training and Self-training code — Compare
DADA-seg (28 rows) MMUDA Towards Robust Semantic Segmentation of Accident Scenes via... code — Compare
DeLiVER (26 rows) CAFuser-CAA CAFuser: Condition-Aware Multimodal Fusion for Robust Semantic... code Syntology ran 1 of 1 samples · 0 unverified Compare
Stanford2D3D Panoramic (25 rows) SFSS-MMSI (RGB+HHA) Single Frame Semantic Segmentation Using Multi-Modal Spherical Images code — Compare
BDD100K val (24 rows) VLTSeg Strong but simple: A Baseline for Domain Generalized Dense... code — Compare
MCubeS (22 rows) StitchFusion (RGB-A-D-N) StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal... code — Compare
CamVid (21 rows) SERNet-Former SERNet-Former: Semantic Segmentation by Efficient Residual Network... code — Compare
COCO-Stuff test (21 rows) VPNeXt VPNeXt -- Rethinking Dense Decoding for Plain Vision Transformer — — Compare
ImageNet-S (20 rows) TEC (ViT-B/16, 224x224, SSL+FT, mmseg) Towards Sustainable Self-supervised Learning code Syntology ran 6 of 8 samples · 2 unverified Compare
ISPRS Potsdam (20 rows) AerialFormer-B AerialFormer: Multi-resolution Transformer for Aerial Image Segmentation code — Compare
LaRS (20 rows) SWIM^2 (Mask2Former) The 2nd Workshop on Maritime Computer Vision (MaCVi) 2024 — — Compare
iSAID (19 rows) SegNeXt-L SegNeXt: Rethinking Convolutional Attention Design for Semantic... code — Compare
LoveDA (19 rows) U-Net (MaxViT-S) U-Net Ensemble for Enhanced Semantic Segmentation in Remote Sensing Imagery — — Compare
KITTI-360 (17 rows) DiPFormer Depth Matters: Exploring Deep Interactions of RGB-D for Semantic... — — Compare
Semantic3D (17 rows) Feature Geometric Net FG-Net: Fast Large-Scale LiDAR Point Clouds Understanding Network... code — Compare
Trans10K (15 rows) Trans4Trans (M) Trans4Trans: Efficient Transformer for Transparent Object... code — Compare
Dark Zurich (14 rows) Refign (HRDA) Refign: Align and Refine for Adaptation of Semantic Segmentation... code Syntology ran 0 of 7 samples · 7 unverified Compare
FMB Dataset (14 rows) RoadFormer+ (RGB-Infrared) RoadFormer+: Delivering RGB-X Scene Parsing through Scale-Aware... — — Compare
UrbanLF (14 rows) CMNeXt (RGB-LF80) Delivering Arbitrary-Modal Semantic Segmentation code Syntology ran 6 of 7 samples · 1 unverified Compare
LIP val (13 rows) Hulk(Finetune, ViT-L) Hulk: A Universal Knowledge Translator for Human-Centric Tasks code Syntology ran 12 of 24 samples · 12 unverified Compare
Nighttime Driving (13 rows) TADP Text-image Alignment for Diffusion-based Perception code — Compare
Vaihingen (13 rows) CMX CMX: Cross-Modal Fusion for RGB-X Semantic Segmentation with Transformers code Syntology ran 2 of 2 samples · 0 unverified Compare
ZJU-RGB-P (13 rows) RoadFormer+ (ConvNeXt-L, RGB-AoLP) RoadFormer+: Delivering RGB-X Scene Parsing through Scale-Aware... — — Compare
EventScape (12 rows) CMX (B4) CMX: Cross-Modal Fusion for RGB-X Semantic Segmentation with Transformers code Syntology ran 2 of 2 samples · 0 unverified Compare
GTAV-to-Cityscapes Labels (12 rows) MIC MIC: Masked Image Consistency for Context-Enhanced Domain Adaptation code — Compare
ISPRS Vaihingen (12 rows) LSKNet-S LSKNet: A Foundation Lightweight Backbone for Remote Sensing code — Compare
ScanNetV2 (12 rows) CMX CMX: Cross-Modal Fusion for RGB-X Semantic Segmentation with Transformers code Syntology ran 2 of 2 samples · 0 unverified Compare
BJRoad (11 rows) CMNeXt Delivering Arbitrary-Modal Semantic Segmentation code Syntology ran 6 of 7 samples · 1 unverified Compare
Potsdam (11 rows) LMFNet-3 LMFNet: An Efficient Multimodal Fusion Approach for Semantic... — — Compare
US3D (11 rows) LMFNet-3 LMFNet: An Efficient Multimodal Fusion Approach for Semantic... — — Compare
Fine-Grained Grass Segmentation Dataset (10 rows) D2LS Dynamic Dictionary Learning for Remote Sensing Image Segmentation code Syntology ran 9 of 12 samples · 3 unverified Compare
SpaceNet 1 (10 rows) MAE+MTP(ViT-L) MTP: Advancing Remote Sensing Foundation Model via Multi-Task Pretraining code Syntology ran 4 of 5 samples · 1 unverified Compare
SYN-UDTIRI (10 rows) RoadFormer+ (B) RoadFormer+: Delivering RGB-X Scene Parsing through Scale-Aware... — — Compare
UAVid (10 rows) U-Net Ensemble U-Net Ensemble for Enhanced Semantic Segmentation in Remote Sensing Imagery — — Compare
COCO (Common Objects in Context) (9 rows) HyperSeg HyperSeg: Towards Universal Visual Segmentation with Large Language Model code Syntology ran 7 of 17 samples · 10 unverified Compare
DDD17 (9 rows) BRENet Rethinking RGB-Event Semantic Segmentation with a Novel... code — Compare
DELIVER (9 rows) CAFuser CAFuser: Condition-Aware Multimodal Fusion for Robust Semantic... code Syntology ran 1 of 1 samples · 0 unverified Compare
DSEC (9 rows) BRENet Rethinking RGB-Event Semantic Segmentation with a Novel... code — Compare
INRIA Aerial Image Labeling (8 rows) UANet(PVT-V2-B2) Building Extraction from Remote Sensing Images via an... code — Compare
LLRGBD-synthetic (8 rows) SMMCL (SegNeXt-B) Understanding Dark Scenes by Contrasting Multi-Modal Observations code — Compare
Mapillary val (8 rows) AO-SegNet Interactive Learning of Intrinsic and Extrinsic Properties for... code — Compare
MCubeS (P) (8 rows) MMSFormer (RGB-A-D) MMSFormer: Multimodal Transformer for Material and Semantic Segmentation code — Compare
SpectralWaste (8 rows) CMX (RGB-HYPER) CMX: Cross-Modal Fusion for RGB-X Semantic Segmentation with Transformers code Syntology ran 2 of 2 samples · 0 unverified Compare
UPLight (8 rows) ShareCMP (B2 RGB-FP) ShareCMP: Polarization-Aware RGB-P Semantic Segmentation code Syntology ran 3 of 3 samples · 0 unverified Compare
FoodSeg103 (7 rows) FoodSAM FoodSAM: Any Food Segmentation code — Compare
KITTI Semantic Segmentation (7 rows) RPVNet [xu2021rpvnet] Spherical Transformer for LiDAR-based 3D Recognition code Syntology ran 4 of 13 samples · 9 unverified Compare
Pothole Mix (7 rows) Baseline - DeepLabv3+ SHREC 2022: pothole and crack detection in the road pavement using... code — Compare
SELMA (7 rows) CMX CMX: Cross-Modal Fusion for RGB-X Semantic Segmentation with Transformers code Syntology ran 2 of 2 samples · 0 unverified Compare
SkyScapes-Dense (7 rows) SkyScapesNet-Dense SkyScapes Fine-Grained Semantic Understanding of Aerial Scenes — — Compare
SYNTHIA-to-Cityscapes (7 rows) HRDA + PiPa PiPa: Pixel- and Patch-wise Self-supervised Learning for Domain... code — Compare
VDD (7 rows) Segformer-B2 VDD: Varied Drone Dataset for Semantic Segmentation code — Compare
ACDC Scribbles (6 rows) ScribFormer ScribFormer: Transformer Makes CNN Work Better for Scribble-based... code — Compare
Event-based Segmentation Dataset (6 rows) Bimodal SegNet Bimodal SegNet: Instance Segmentation Fusing Events and RGB Frames... code — Compare
GAMUS (6 rows) TIMF GAMUS: A Geometry-aware Multi-modal Semantic Segmentation... code — Compare
Porto (6 rows) CMNeXt Delivering Arbitrary-Modal Semantic Segmentation code Syntology ran 6 of 7 samples · 1 unverified Compare
Stanford2D3D - RGBD (6 rows) CMX (SegFormer-B4) CMX: Cross-Modal Fusion for RGB-X Semantic Segmentation with Transformers code Syntology ran 2 of 2 samples · 0 unverified Compare
SynPASS (6 rows) Trans4PASS+ Behind Every Domain There is a Shift: Adapting Distortion-aware... code — Compare
TLCGIS (6 rows) SA-Gate Bi-directional Cross-Modality Feature Propagation with... code — Compare
FLAIR (French Land cover from Aerospace ImageRy) (5 rows) Ensemble-04 MiT-0 MiT-1 RNX-1 RNX-2 Modernized Training of U-Net for Aerial Semantic Segmentation code — Compare
Hypersim (5 rows) EMSANet (2x ResNet-34 NBt1D) PanopticNDT: Efficient and Robust Panoptic Mapping code Syntology ran 0 of 12 samples · 12 unverified Compare
RELLIS-3D Dataset (5 rows) Swiftnet Semantic Segmentation with High Inference Speed in Off-Road Environments code — Compare
Replica (5 rows) LabelMaker LABELMAKER: Automatic Semantic Label Generation from RGB-D Trajectories code Syntology ran 3 of 3 samples · 0 unverified Compare
ShapeNet (5 rows) PatchFormer PatchFormer: An Efficient Point Transformer with Patch Attention — — Compare
Synthetic Bathing Perception (5 rows) CMX-SRA CMX: Cross-Modal Fusion for RGB-X Semantic Segmentation with Transformers code Syntology ran 2 of 2 samples · 0 unverified Compare
Toronto-3D L002 (5 rows) EyeNet Human Vision Based 3D Point Cloud Semantic Segmentation of... code — Compare
BIG (4 rows) PSPNet + CascadePSP CascadePSP: Toward Class-Agnostic and Very High-Resolution... code Syntology ran 2 of 7 samples · 5 unverified Compare
CC3M-TagMask (4 rows) TTD (TCL) TTD: Text-Tag Self-Distillation Enhancing Image-Text Alignment in... code Syntology ran 1 of 1 samples · 0 unverified Compare
Fine-Grained Cloud Segmentation Dataset (4 rows) D2LS Dynamic Dictionary Learning for Remote Sensing Image Segmentation code Syntology ran 9 of 12 samples · 3 unverified Compare
Lombardia Sentinel-2 Image Time Series for Crop Mapping (4 rows) UNet3D Enhancing crop segmentation in satellite image time-series with... code — Compare
Matterport3D (4 rows) SFSS-MMSI (RGB+Depth) Single Frame Semantic Segmentation Using Multi-Modal Spherical Images code — Compare
Okutama Drone and Swiss Drone Dataset (4 rows) DeepLabv3+‐ResNet‐101 Deep learning with RGB and thermal images onboard a drone for... — — Compare
PETRAW (4 rows) NCC Next PEg TRAnsfer Workflow recognition challenge report: Does... — — Compare
Structured3D (4 rows) SFSS-MMSI (RGB+Depth+Normal) Single Frame Semantic Segmentation Using Multi-Modal Spherical Images code — Compare
THUD Robotic Dataset (4 rows) SA-Gate Bi-directional Cross-Modality Feature Propagation with... code — Compare
CEMS-W (3 rows) UPerNet (RN50) Robust Burned Area Delineation through Multitask Learning code — Compare
dacl10k v1 testdev (3 rows) FPN EfficientNet-B4 w/ Aux loss dacl10k: Benchmark for Semantic Bridge Damage Segmentation code — Compare
DIVA-HisDB (3 rows) U-Net DIVA-DAF: A Deep Learning Framework for Historical Document Image Analysis code — Compare
Montgomery County X-ray Set (3 rows) UNETR + SS-CXR SPCXR: Self-supervised Pretraining using Chest X-rays Towards a... — — Compare
PASCAL VOC 2011 test (3 rows) Plugin network Plugin Networks for Inference under Partial Evidence code — Compare
PASTIS (3 rows) Exchanger+Mask2Former Revisiting the Encoding of Satellite Image Time Series code — Compare
Potsdam (3 rows) HRNet-48 Deep High-Resolution Representation Learning for Visual Recognition code Syntology ran 3 of 34 samples · 31 unverified Compare
SIFT-flow (3 rows) RBE2E Region-based semantic segmentation with end-to-end training code — Compare
Stanford2D3D Panoramic - RGBD (3 rows) CBFC Complementary Bi-directional Feature Compression for Indoor 360°... — — Compare
SYNTHIA (3 rows) CGA-Net CGA-Net: Category Guided Aggregation for Point Cloud Semantic Segmentation code — Compare
US3D (3 rows) DeepLabV3+ Encoder-Decoder with Atrous Separable Convolution for Semantic... code Syntology ran 43 of 72 samples · 29 unverified Compare
38-Cloud (2 rows) Cloud-Net+ Cloud and Cloud Shadow Segmentation for Remote Sensing Imagery via... code — Compare
AI-TOD (2 rows) Unet++(ResNet-50) UNet++: A Nested U-Net Architecture for Medical Image Segmentation code Syntology ran 5 of 28 samples · 23 unverified Compare
ApolloScape (2 rows) ERFNet-IntRA-KD (ours) Inter-Region Affinity Distillation for Road Marking Segmentation code Syntology ran 0 of 4 samples · 4 unverified Compare
Cityscapes (2 rows) SPFNet34M S²-FPN: Scale-ware Strip Attention Guided Feature Pyramid Network... code — Compare
Cleargrasp (Novel) (2 rows) Cleargrasp ClearGrasp: 3D Shape Estimation of Transparent Objects for Manipulation code Syntology ran 0 of 8 samples · 8 unverified Compare
Endoscapes (2 rows) MoCo V2 Surg SSL - DeepLabv3+ head Dissecting Self-Supervised Learning Methods for Surgical Computer Vision code — Compare
Freiburg Forest (2 rows) SSMA Self-Supervised Model Adaptation for Multimodal Semantic Segmentation code — Compare
Graz-02 (2 rows) VOLO-D5 VOLO: Vision Outlooker for Visual Recognition code Syntology ran 1 of 6 samples · 5 unverified Compare
HePIC 🏛️ (2 rows) BIM-Net++ Fully Automated Scan-to-BIM Via Point Cloud Instance Segmentation code — Compare
HERA RFI Detection (2 rows) Nearest Latent Neighbours RFI Detection with Spiking Neural Networks code — Compare
Kvasir-Instrument (2 rows) DoubleUNet DoubleU-Net: A Deep Convolutional Neural Network for Medical Image... code Syntology ran 2 of 5 samples · 3 unverified Compare
LOFAR RFI Detection (2 rows) Nearest Latent Neighbours RFI Detection with Spiking Neural Networks code — Compare
MUSES: MUlti-SEnsor Semantic perception dataset (2 rows) CAFuser (Swin-T) CAFuser: Condition-Aware Multimodal Fusion for Robust Semantic... code Syntology ran 1 of 1 samples · 0 unverified Compare
PASCAL VOC 2007 (2 rows) GALDNet Global Aggregation then Local Distribution in Fully Convolutional Networks code Syntology ran 2 of 4 samples · 2 unverified Compare
PH2 (2 rows) MobileUNETR MobileUNETR: A Lightweight End-To-End Hybrid Vision Transformer... code — Compare
SkyScapes-Lane (2 rows) SkyScapesNet-Lane SkyScapes Fine-Grained Semantic Understanding of Aerial Scenes — — Compare
SYNTHIA-CVPR’16 (2 rows) SSMA Self-Supervised Model Adaptation for Multimodal Semantic Segmentation code — Compare
AIRS (1 row) ICT-Net Semantic Segmentation from Remote Sensor Data and the Exploitation... code — Compare
ARCH2S (1 row) BIM-Net Fully Automated Scan-to-BIM Via Point Cloud Instance Segmentation code — Compare
ATLANTIS (1 row) Erfani et al. ATLANTIS: A Benchmark for Semantic Segmentation of Waterbody Images code — Compare
BDD (1 row) FasterSeg FasterSeg: Searching for Faster Real-time Semantic Segmentation code Syntology ran 2 of 13 samples · 11 unverified Compare
Cam2BEV (1 row) uNetXST A Sim2Real Deep Learning Approach for the Transformation of Images... code Syntology ran 0 of 6 samples · 6 unverified Compare
Cityscapes 3D (1 row) TaskPrompter Joint 2D-3D Multi-Task Learning on Cityscapes-3D: 3D Detection,... code — Compare
Cityscapes VIPriors subset (1 row) EfficientSeg EfficientSeg: An Efficient Semantic Segmentation Network code — Compare
COCO-Stuff (1 row) Deeplab v2 COCO-Stuff: Thing and Stuff Classes in Context code Syntology ran 1 of 29 samples · 28 unverified Compare
COCO-Stuff-27 (1 row) DiffSeg (512) Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation... code Syntology ran 3 of 3 samples · 0 unverified Compare
COCO-Stuff full (1 row) SegFormer-B5 (Single Scale) SegFormer: Simple and Efficient Design for Semantic Segmentation... code Syntology ran 48 of 86 samples · 38 unverified Compare
dacl10k v1 testfinal (1 row) FPN EfficientNet-B4 dacl10k: Benchmark for Semantic Bridge Damage Segmentation code — Compare
DeLiVER test (1 row) CAFuser CAFuser: Condition-Aware Multimodal Fusion for Robust Semantic... code Syntology ran 1 of 1 samples · 0 unverified Compare
DroneDeploy (1 row) DLv3+ (Xception65) Aerial Imagery Pixel-level Segmentation code — Compare
Forward-Looking Sonar Marine Debris Datasets (1 row) Unet+RN34 The Marine Debris Dataset for Forward-Looking Sonar Semantic Segmentation code — Compare
FP4S (1 row) FP4S Floor Plan Image Segmentation Via Scribble-Based Semi-Weakly... code — Compare
HAM10000 (1 row) MFSNet MFSNet: A Multi Focus Segmentation Network for Skin Lesion Segmentation code — Compare
ISIC 2017 (1 row) MFSNet MFSNet: A Multi Focus Segmentation Network for Skin Lesion Segmentation code — Compare
LandCover.ai (1 row) U-Net (ConvFormer-M36) U-Net Ensemble for Enhanced Semantic Segmentation in Remote Sensing Imagery — — Compare
ManipalUAVid (1 row) UVid-Net UVid-Net: Enhanced Semantic Segmentation of UAV Aerial Videos by... code — Compare
Mila Simulated Floods (1 row) FloodTransformer (Ours) Transformer-based Flood Scene Segmentation for Developing Countries — — Compare
MixedWM38 (1 row) WaferSegClassNet WaferSegClassNet -- A Light-weight Network for Classification and... code — Compare
OpenEDS (1 row) RITnet RITnet: Real-time Semantic Segmentation of the Eye for Gaze Tracking code — Compare
PASCAL VOC (1 row) SegCLIP SegCLIP: Patch Aggregation with Learnable Centers for... code — Compare
PASCAL VOC 2010 test (1 row) SIW Scaling up Multi-domain Semantic Segmentation with Sentence Embeddings — — Compare
PASCAL VOC 2011 (1 row) DLDL-8s+CRF Deep Label Distribution Learning with Label Ambiguity code — Compare
PASCAL VOC 2012 (1 row) DLDL-8s+CRF Deep Label Distribution Learning with Label Ambiguity code — Compare
PASTIS-R (1 row) Late Fusion Multi-Modal Temporal Attention Models for Crop Mapping from... code — Compare
RUGD (1 row) GA-Nav GANav: Efficient Terrain Segmentation for Robot Navigation in... code — Compare
SBCoseg (1 row) Dice loss + IS-Triplet loss Improving Image co-segmentation via Deep Metric Learning — — Compare
SemanticPOSS (1 row) TFNet TFNet: Exploiting Temporal Cues for Fast and Accurate LiDAR... — — Compare
STARE (1 row) UNet U-Net: Convolutional Networks for Biomedical Image Segmentation code Syntology ran 510 of 757 samples · 247 unverified Compare
SWIMSEG (1 row) ACLNet ACLNet: An Attention and Clustering-based Cloud Segmentation Network code — Compare
SWINSEG (1 row) ACLNet ACLNet: An Attention and Clustering-based Cloud Segmentation Network code — Compare
SWINySEG (1 row) ACLNet ACLNet: An Attention and Clustering-based Cloud Segmentation Network code — Compare
UTFPR-SBD3 (1 row) EPYNET EPYNET: Efficient Pyramidal Network for Clothing Segmentation — — Compare
WildDash (1 row) SIW Scaling up Multi-domain Semantic Segmentation with Sentence Embeddings — — Compare
ADE20K-150 (0 rows) no rows in the archive — —

Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

347 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 347 until expanded.

COCO (Common Objects in Context)CityscapesKITTIShapeNetScanNetADE20KNYUv2DAVISEuroSATSYNTHIAS3DISSUN RGB-DBDD100KMatterport3DRefCOCOReplicaGTA5COCO-StuffPASCAL ContextDAVIS 2017KITTI-360CamVidVisDA-2017HAM10000HelenKvasir-SEGPASCAL VOCSUNCGLabelMePASCAL-5iYouTube-VIS 2019Objects365PartNet2D-3D-SSTAREVirtual KITTIMake3DPASCAL VOC 2007SUN3DKvasirPASCAL VOC 2012 testHypersimSegTrack-v2AVEMapillary Vistas DatasetIDDMedical Segmentation DecathlonBigEarthNetStructured3DPROMISE12iSAIDLoveDAReferItGameApolloScapeSemanticPOSSBraTS 2015CoNSePSemantic3DPlaces365LIPPanNukeDark ZurichLost and FoundRELLIS-3DWoodScapeUAVidVirtual KITTI 2xBD3D-FUTUREWildDashStanford BackgroundSynscapes3RScanDensePASSImageNet-SDDD17KITTI RoadRoadTracerSlakh2100WORDOpen Images V4STPLS3DPST900OxUvaSEN12MSSUIMTrans10KOpenEDSPhraseCutDeepFashion2InteriorNetOCIDPartImageNetSegTHORSensatUrbanGIDISIC 2018 Task 1ISPRS PotsdamNighttime DrivingScribbleSupIKEA ASMROSEDADA-segToronto-3DDeepWeedsDigestPathDSECHOI4DCARRADASatlasINRIA Aerial Image LabelingMSegAVSBenchFMB DatasetPASCAL VOC 2011DeepFishDELIVERFoodSeg103Hi4DISPRS VaihingenKvasir-InstrumentMCubeSModaNetAgriculture-VisionAI-TODCrossMoDAPASTISMoNuSegSceneNetCryoNuSegDADA-2000DIVA-HisDBParis-Lille-3DRoadAnomaly21FashionpediaMHPREFUGE ChallengeBraTS 2016Cityscapes 3DHM3DSemTTPLAWaterScenesPartial-iLIDSSemanticSTFTrashCanWildScenesLandCover.aiRailSem19SD-198So2Sat LCZ42TACO2-PM Vessel DatasetACDC ScribblesCropAndWeedFine-Grained Grass Segmentation DatasetPH2SpaceNet 2AIRSBIGBIMCV COVID-19EgoHOSEntitySegPerSegPhenoBenchSAMRSSketchySceneTSSBUIISDTTD-MobileEchonet-DynamicMontgomery County X-ray SetRIT-18SynthCityZJU-RGB-PNDD20OpenEDS2020OST300Plant Seedlings DatasetSunnybrook Cardiac DataEarthVQAFreiburg ForestGFFMatterportLayoutOpen Images V7SWIMSEGVDDXImageNet-1225kTreesBraTS 2014BUP20Cata7CC3M-TagMaskGolfDBIDDAISBDAKvasir-Sessile datasetLabPicsLaRSm2caiSegMarket1501-AttributesNERDS 360Rent3DSpaceNet 1SWINSEGTAS500UPLightWGISDZeroWasteAeroRITCrackVision12KDukeMTMC-attributeEDENEndoscapesFine-Grained Cloud Segmentation DatasetFLAIR (French Land cover from Aerospace ImageRy)MCubeS (P)Middlebury 2001MJU-WasteMUADMUSES: MUlti-SEnsor Semantic perception datasetPASTIS-RSWINySEGSwiss3DCitiesUSIS10KVLOG DatasetWISDOMAircraft Context DatasetAll-day CityScapesARCH2SBRIGHTCheXlocalizeDeepSportRadar-v1DOORSFive-Billion-PixelsHelixNetHephaestusImage and Video AdvertisementsKvasirCapsule-SEGMedico automatic polyp segmentation challenge (dataset)OAM-TCDOpenTTGamesPETRAWPICRobotPushRuralscapesSpaceNet MVOIUNDDVocalFoldsATLANTISCalCROP21Cam2BEVCityscapes VIPriors subsetColorectal AdenomaCumulodacl10kEMDS-6Endotect Polyp Segmentation Challenge DatasetExtended Agriculture-VisionForward-Looking Sonar Marine Debris DatasetsFPDSFracAtlasG-VUEHERA RFI DetectionHuTicsLOFAR RFI DetectionLombardia Sentinel-2 Image Time Series for Crop MappingMagicBathyNetMila Simulated FloodsMUSICMVTec D2SNVGazeOADATODMSOmniCityPAX-Ray++Photi-LakeIcePothole MixPotsdamRetinal MicrosurgeryRoad Scene GraphRUGDSB20SegRCDBSemanticSpray DatasetTAS-NIRTopology Optimization DatasetU-DIADS-BibVistas-NPBCSDBurned Area Delineation from Satellite ImageryCARLA2RealCCSECEMS-WCiona17CongNaMulCUTSDADEDroneDeployEBHI-SegFP4SFreiburg TerrainsGUISS datasetHASCDHePIC 🏛️HuTu 80IndraEyeIntegrated Drone Dataset for Semantic SegmentationLIB-HSIMeta OmniumMixedWM38MLSMOSADMTNeuroMulti-species fruit flower detection datasetsMultispectral and HD vineyard orthomosaics from central PortugalOkutama Drone and Swiss Drone DatasetPanoramic Video Panoptic Segmentation DatasetPolyp ASHRisk-Aware Planning DatasetSBCosegSemantic Segmentation Vineyard RowsSemanticSugarBeetsSemanticUSLSen4AgriNetSYNTHIA-PANOThe RobotriXThermal Face DatabaseTimberVisionTintoTransProteusUncertainty Quantification for Underwater Object SegmentationUTFPR-SBD3VizWiz-FewShotXKCDColorsAachen-Heerlen Annotated Steel Microstructure DatasetCorn Seeds DatasetDMSDronescapesLemons quality control datasetMVP-24KOCT5kOpenSurfacesOUTBACK: A Multimodal Synthetic Dataset for Rural Australian Off-road Robot NavigationPAXRayPSIEUAV-based multispectral vineyardsWTA/TLA

Subtasks archive 2025-07-28

29 subtasks in the archive's task tree.

Parent tasks archive 2025-07-28

Most implemented papers archive 2025-07-28

30 shown of 6,644 papers with code (14,763 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

  • 18 May 2015 487 repositories listed Syntology ran 510 of 757 samples · 247 unverified · 426 pointer-only (licence)
    There is large consent that successful training of deep networks requires many thousand annotated training samples.
  • 10 Dec 2015 484 repositories listed Syntology ran 230 of 377 samples · 147 unverified · 187 pointer-only (licence)
    Deep residual nets are foundations of our submissions to ILSVRC & COCO 2015 competitions, where we also won the 1st places on the tasks of ImageNet detection, ImageNet localization, COCO detection, and COCO segmentation.
  • 20 Mar 2017 179 repositories listed Syntology ran 42 of 140 samples · 98 unverified · 23 pointer-only (licence)
    Our approach efficiently detects objects in an image while simultaneously generating a high-quality segmentation mask for each instance.
  • 13 Jan 2018 159 repositories listed Syntology ran 85 of 111 samples · 26 unverified · 64 pointer-only (licence)
    In this paper we describe a new mobile architecture, MobileNetV2, that improves the state of the art performance of mobile models on multiple tasks and benchmarks as well as across a spectrum of different model sizes.
  • 22 Oct 2020 158 repositories listed Syntology ran 281 of 419 samples · 138 unverified · 154 pointer-only (licence)
    While the Transformer architecture has become the de-facto standard for natural language processing tasks, its applications to computer vision remain limited.
  • 17 Jun 2019 142 repositories listed Syntology ran 14 of 82 samples · 68 unverified
    In this paper, we introduce the various features of this toolbox.
  • 2 Dec 2016 110 repositories listed Syntology ran 89 of 164 samples · 75 unverified · 90 pointer-only (licence)
    Point cloud is an important type of geometric data structure.
  • 2 Apr 2019 87 repositories listed Syntology ran 13 of 40 samples · 27 unverified · 18 pointer-only (licence)
    By eliminating the predefined set of anchor boxes, FCOS completely avoids the complicated computation related to anchor boxes such as calculating overlapping during training.
  • 9 Dec 2016 85 repositories listed Syntology ran 16 of 51 samples · 35 unverified · 11 pointer-only (licence)
    Feature pyramids are a basic component in recognition systems for detecting objects at different scales.
  • 25 Mar 2021 80 repositories listed Syntology ran 108 of 207 samples · 99 unverified · 43 pointer-only (licence)
    This paper presents a new vision Transformer, called Swin Transformer, that capably serves as a general-purpose backbone for computer vision.
  • 7 Feb 2018 78 repositories listed Syntology ran 43 of 72 samples · 29 unverified · 40 pointer-only (licence)
    The former networks are able to encode multi-scale contextual information by probing the incoming features with filters or pooling operations at multiple rates and multiple effective fields-of-view, while the latter…
  • 17 Jun 2017 77 repositories listed Syntology ran 3 of 7 samples · 4 unverified · 3 pointer-only (licence)
    To handle the problem of segmenting objects at multiple scales, we design modules which employ atrous convolution in cascade or in parallel to capture multi-scale context by adopting multiple atrous rates.
  • 2 Nov 2015 74 repositories listed Syntology ran 9 of 44 samples · 35 unverified · 10 pointer-only (licence)
    We show that SegNet provides good performance with competitive inference time and more efficient inference memory-wise as compared to other architectures.
  • 7 Jun 2017 68 repositories listed Syntology ran 36 of 67 samples · 31 unverified · 26 pointer-only (licence)
    By exploiting metric space distances, our network is able to learn local features with increasing contextual scales.
  • 6 May 2019 67 repositories listed Syntology ran 58 of 105 samples · 47 unverified · 46 pointer-only (licence)
    We achieve new state of the art results for mobile classification, detection and segmentation.
  • 4 Dec 2016 67 repositories listed Syntology ran 7 of 29 samples · 22 unverified · 5 pointer-only (licence)
    Scene parsing is challenging for unrestricted open vocabulary and diverse scenes.
  • 11 Nov 2021 58 repositories listed Syntology ran 71 of 137 samples · 66 unverified · 73 pointer-only (licence)
    Our MAE approach is simple: we mask random patches of the input image and reconstruct the missing pixels.
  • 10 Jan 2022 54 repositories listed Syntology ran 54 of 80 samples · 26 unverified · 11 pointer-only (licence)
    The "Roaring 20s" of visual recognition began with the introduction of Vision Transformers (ViTs), which quickly superseded ConvNets as the state-of-the-art image classification model.
  • 14 Nov 2014 51 repositories listed Syntology ran 3 of 4 samples · 1 unverified · 4 pointer-only (licence)
    Convolutional networks are powerful visual models that yield hierarchies of features.
  • 7 Jun 2016 49 repositories listed Syntology ran 5 of 30 samples · 25 unverified
    The ability to perform pixel-wise semantic segmentation in real-time is of paramount importance in mobile applications.
  • 4 Apr 2019 48 repositories listed Syntology ran 10 of 21 samples · 11 unverified · 6 pointer-only (licence)
    Then we produce instance masks by linearly combining the prototypes with the mask coefficients.
  • 2 Jun 2016 47 repositories listed Syntology ran 28 of 63 samples · 35 unverified · 16 pointer-only (licence)
    ASPP probes an incoming convolutional feature layer with filters at multiple sampling rates and effective fields-of-views, thus capturing objects as well as image context at multiple scales.
  • 20 Aug 2019 42 repositories listed Syntology ran 3 of 34 samples · 31 unverified · 16 pointer-only (licence)
    High-resolution representations are essential for position-sensitive vision problems, such as human pose estimation, semantic segmentation, and object detection.
  • 9 Apr 2019 39 repositories listed Syntology ran 3 of 18 samples · 15 unverified · 5 pointer-only (licence)
    The proposed approach achieves superior results to existing single-model networks on COCO object detection.
  • 17 Mar 2017 38 repositories listed Syntology ran 3 of 9 samples · 6 unverified · 3 pointer-only (licence)
    Convolutional neural networks (CNNs) are inherently limited to model geometric transformations due to the fixed geometric structures in its building modules.
  • 1 May 2014 38 repositories listed Syntology ran 1 of 7 samples · 6 unverified
    We present a new dataset with the goal of advancing the state-of-the-art in object recognition by placing the question of object recognition in the context of the broader question of scene understanding.
  • 11 Apr 2018 37 repositories listed Syntology ran 8 of 28 samples · 20 unverified · 7 pointer-only (licence)
    We propose a novel attention gate (AG) model for medical imaging that automatically learns to focus on target structures of varying shapes and sizes.
  • 20 May 2016 37 repositories listed
    Convolutional networks are powerful visual models that yield hierarchies of features.
  • 19 Apr 2020 36 repositories listed Syntology ran 8 of 48 samples · 40 unverified · 23 pointer-only (licence)
    It is well known that featuremap attention and multi-path representation are important for visual recognition.
  • 3 Dec 2019 36 repositories listed Syntology ran 11 of 43 samples · 32 unverified
    Then we produce instance masks by linearly combining the prototypes with the mask coefficients.

Syntology lines on 29 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections