Papers › Strong but simple: A Baseline for Domain Generalized Dense Perception by CLIP-based...

Strong but simple: A Baseline for Domain Generalized Dense Perception by CLIP-based Transfer Learning

4 Dec 2023arXiv:2312.02021archive 2025-07-28

Christoph Hümmer, Manuel Schwonberg, Liangwei Zhou, Hu Cao, Alois Knoll, Hanno Gottschalk

Domain generalization (DG) remains a significant challenge for perception based on deep neural networks (DNNs), where domain shifts occur due to synthetic data, lighting, weather, or location changes. Vision-language models (VLMs) marked a large step for the generalization capabilities and have been already applied to various tasks. Very recently, first approaches utilized VLMs for domain generalized segmentation and object detection and obtained strong generalization. However, all these approaches rely on complex modules, feature augmentation frameworks or additional models. Surprisingly and in contrast to that, we found that simple fine-tuning of vision-language pre-trained models yields competitive or even stronger generalization results while being extremely simple to apply. Moreover, we found that vision-language pre-training consistently provides better generalization than the previous standard of vision-only pre-training. This challenges the standard of using ImageNet-based transfer learning for domain generalization. Fully fine-tuning a vision-language pre-trained model is capable of reaching the domain generalization SOTA when training on the synthetic GTA5 dataset. Moreover, we confirm this observation for object detection on a novel synthetic-to-real benchmark. We further obtain superior generalization capabilities by reaching 77.9% mIoU on the popular Cityscapes-to-ACDC benchmark. We also found improved in-domain generalization, leading to an improved SOTA of 86.4% mIoU on the Cityscapes test set marking the first place on the leaderboard.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

VLTSeg/VLTSeg mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Domain GeneralizationObject DetectionRobust Object DetectionSegmentationSemantic SegmentationTransfer LearningVision-Language Segmentationobject-detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Domain Generalization GTA-to-Avg(Cityscapes,BDD,Mapillary) VLTSeg mIoU 63.5 #6 of 24 Archive leaderboard report
Domain Generalization GTA5-to-Cityscapes VLTSeg (EVA02-CLIP-L) mIoU 65.6 #4 of 8 Archive leaderboard report
Robust Object Detection DWD VLTDet mPC [AP50] 36.9 #4 of 12 Archive leaderboard report
Semantic Segmentation BDD100K val VLTSeg mIoU 72.5 #1 of 24 Archive leaderboard report
Semantic Segmentation Cityscapes test VLTSeg Mean IoU (class) 86.4 #1 of 105 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

CLIPSET

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections