Home › Datasets › task › Domain Generalization

Domain Generalization datasets

archive 2025-07-28

31 datasets carry the task tag "Domain Generalization" (the task itself: Domain Generalization), ordered by the archive's paper count. Page 1 of 1: 31 shown of 31. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Domain Generalization datasets 1–31 of 31

Cityscapes is a large-scale database which focuses on semantic understanding of urban street scenes.
3,702 papers · 51 benchmarks
Fashion-MNIST is a dataset comprising of 28×28 grayscale images of 70,000 fashion products from 10 categories, with 7,000 images per category.
3,202 papers · 15 benchmarks
Office-Home is a benchmark dataset for domain adaptation which contains 4 domains where each domain consists of 65 categories.
1,074 papers · 11 benchmarks
PACS (Photo-Art-Cartoon-Sketch)
PACS is an image dataset for domain generalization.
668 papers · 10 benchmarks
ImageNet-C is an open source data set that consists of algorithmically generated corruptions (blur, noise) applied to the ImageNet test-set.
602 papers · 4 benchmarks
Common corruptions dataset for CIFAR10
494 papers · 2 benchmarks
ImageNet-R (ImageNet-Rendition)
ImageNet-R(endition) contains art, cartoons, deviantart, graffiti, embroidery, graphics, origami, paintings, patterns, plastic objects, plush objects, sculptures, sketches, tattoos, toys, and video game renditions of ImageNet classes.
481 papers · 5 benchmarks
The ImageNet-A dataset consists of real-world, unmodified, and naturally occurring examples that are misclassified by ResNet models.
431 papers · 5 benchmarks
ImageNet-Sketch data set consists of 50,889 images, approximately 50 images for each of the 1000 ImageNet classes.
268 papers · 3 benchmarks
The 'shape bias' dataset was introduced in Geirhos et al.
120 papers · 1 benchmark
The Stylized-ImageNet dataset is created by removing local texture cues in ImageNet while retaining global shape information on natural images via AdaIN style transfer.
106 papers · 1 benchmark
WildDash is a benchmark evaluation method is presented that uses the meta-information to calculate the robustness of a given algorithm with respect to the individual hazards.
47 papers · 2 benchmarks
VLCS is a dataset to test for domain generalization.
33 papers · 1 benchmark
ImageNet-P consists of noise, blur, weather, and digital distortions.
32 papers · 1 benchmark
The goal of NICO Challenge is to facilitate the OOD (Out-of-Distribution) generalization in visual recognition through promoting the research on the intrinsic learning mechanisms with native invariance and generalization ability.
32 papers · 1 benchmark
We introduce ACDC, the Adverse Conditions Dataset with Correspondences for training and testing semantic segmentation methods on adverse visual conditions.
31 papers · 5 benchmarks
Our goal is to improve upon the status quo for designing image classification models trained in one domain that perform well on images from another domain.
22 papers · 3 benchmarks
The MSK dataset is a dataset for lesion recognition from the Memorial Sloan-Kettering Cancer Center.
15 papers · 0 benchmarks
Super-CLEVR is a dataset for Visual Question Answering (VQA) where different factors in VQA domain shifts can be isolated in order that their effects can be studied independently.
13 papers · 0 benchmarks
NICO (Non-I.I.D. Image dataset with Contexts)
I.I.D.
9 papers · 2 benchmarks
Wild-Time is a benchmark of 5 datasets that reflect temporal distribution shifts arising in a variety of real-world applications, including patient prognosis and news classification.
6 papers · 0 benchmarks
The Sims4Action Dataset: a videogame-based dataset for Synthetic→Real domain adaptation for human activity recognition.
5 papers · 0 benchmarks
As part of an ongoing worldwide effort to comprehend and monitor insect biodiversity, we present the BIOSCAN-5M Insect dataset to the machine learning community.
4 papers · 0 benchmarks
Human-Animal-Cartoon (HAC) dataset consists of seven actions (‘sleeping’, ‘watching tv’, ‘eating’, ‘drinking’, ‘swimming’, ‘running’, and ‘opening door’) performed by humans, animals, and cartoon figures, forming three different domains.
4 papers · 0 benchmarks
We design an all-day semantic segmentation benchmark all-day CityScapes.
3 papers · 1 benchmark
CFC (Caltech Fish Counting Dataset)
Caltech Fish Counting Dataset (CFC) is a large-scale dataset for detecting, tracking, and counting fish in sonar videos.
3 papers · 0 benchmarks
This prostate MRI segmentation dataset is collected from six different data sources.
3 papers · 0 benchmarks
In our benchmark WHYSHIFT, we explore distribution shifts on 5 real-world tabular datasets from the economic and traffic sectors with natural spatiotemporal distribution shifts.We only pick 7 typical settings out of 22 settings and select…
3 papers · 0 benchmarks
Domain-independent anomalies datasets (Domain-independent anomalies datasets (adaptions of the MVTec Anomaly Detection dataset))
An adaption of the MVTec Anomaly Detection dataset, presented in the paper "Domain-independent detection of known anomalies".
1 paper · 1 benchmark
fNIRS2MW (The Tufts fNIRS to Mental Workload Dataset)
The Tufts fNIRS to Mental Workload (fNIRS2MW) open-access dataset is a new dataset for building machine learning classifiers that can consume a short window (30 seconds) of multivariate fNIRS recordings and predict the mental workload…
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.