Home › Datasets › task › Domain Adaptation

Domain Adaptation datasets

archive 2025-07-28

96 datasets carry the task tag "Domain Adaptation" (the task itself: Domain Adaptation), ordered by the archive's paper count. Page 1 of 2: 48 shown of 96. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Domain Adaptation datasets 1–48 of 96

The MNIST database (Modified National Institute of Standards and Technology database) is a large collection of handwritten digits.
7,651 papers · 44 benchmarks
SVHN (Street View House Numbers)
Street View House Numbers (SVHN) is a digit classification benchmark dataset that contains 600,000 32×32 RGB images of printed digits (from 0 to 9) cropped from pictures of house number plates.
3,406 papers · 12 benchmarks
Office-Home is a benchmark dataset for domain adaptation which contains 4 domains where each domain consists of 65 categories.
1,074 papers · 11 benchmarks
PACS (Photo-Art-Cartoon-Sketch)
PACS is an image dataset for domain generalization.
668 papers · 10 benchmarks
Office-31 (Office Dataset)
The Office dataset contains 31 object categories in three domains: Amazon, DSLR and Webcam.
643 papers · 7 benchmarks
SYNTHIA (SYNTHetic Collection of Imagery and Annotations)
The SYNTHIA dataset is a synthetic dataset that consists of 9400 multi-viewpoint photo-realistic frames rendered from a virtual city and comes with pixel-level semantic annotations for 13 classes.
538 papers · 10 benchmarks
Datasets drive vision progress, yet existing driving datasets are impoverished in terms of visual content and supported tasks to study multitask learning for autonomous driving.
469 papers · 16 benchmarks
USPS is a digit dataset automatically scanned from envelopes by the U.S.
459 papers · 2 benchmarks
The Replica Dataset is a dataset of high quality reconstructions of a variety of indoor spaces.
414 papers · 4 benchmarks
GTA5 (Grand Theft Auto 5)
The GTA5 dataset contains 24966 synthetic images with pixel level semantic annotation.
412 papers · 7 benchmarks
GTSRB (German Traffic Sign Recognition Benchmark)
The German Traffic Sign Recognition Benchmark (GTSRB) contains 43 classes of traffic signs, split into 39,209 training images and 12,630 test images.
374 papers · 5 benchmarks
ImageNet-Sketch data set consists of 50,889 images, approximately 50 images for each of the 1000 ImageNet classes.
268 papers · 3 benchmarks
Foggy Cityscapes is a synthetic foggy dataset which simulates fog on real scenes.
249 papers · 7 benchmarks
VisDA-2017 is a simulation-to-real dataset for domain adaptation with over 280,000 images across 12 categories in the training, validation and testing domains.
223 papers · 6 benchmarks
OpenSubtitles is collection of multilingual parallel corpora.
214 papers · 3 benchmarks
MNIST-M is created by combining MNIST digits with the patches randomly extracted from color photos of BSDS500 as their background.
193 papers · 1 benchmark
FSD50K (Freesound Database 50K)
Freesound Dataset 50k (or FSD50K for short) is an open dataset of human-labeled sound events containing 51,197 Freesound clips unequally distributed in 200 classes drawn from the AudioSet Ontology.
155 papers · 2 benchmarks
IDD (Indian Driving Dataset)
IDD is a dataset for road scene understanding in unstructured environments used for semantic segmentation and object detection for autonomous driving.
98 papers · 1 benchmark
The ImageCLEF-DA dataset is a benchmark dataset for ImageCLEF 2014 domain adaptation challenge, which contains three domains: Caltech-256 (C), ImageNet ILSVRC 2012 (I) and Pascal VOC 2012 (P).
96 papers · 1 benchmark
The VGG Face dataset is face identity recognition dataset that consists of 2,622 identities.
94 papers · 0 benchmarks
SIM10k is a synthetic dataset containing 10,000 images, which is rendered from the video game Grand Theft Auto V (GTA5).
92 papers · 3 benchmarks
ASPEC (Asian Scientific Paper Excerpt Corpus)
ASPEC, Asian Scientific Paper Excerpt Corpus, is constructed by the Japan Science and Technology Agency (JST) in collaboration with the National Institute of Information and Communications Technology (NICT).
87 papers · 0 benchmarks
LoveDA (Remote Sensing Land-Cover Dataset for Domain Adaptive Semantic Segmentation)
1.
81 papers · 1 benchmark
The Machine Translation of Noisy Text (MTNT) dataset is a Machine Translation dataset that consists of noisy comments on Reddit and professionally sourced translation.
52 papers · 0 benchmarks
Synscapes is a synthetic dataset for street scene parsing created using photorealistic rendering techniques, and show state-of-the-art results for training and validation as well as new types of analysis.
46 papers · 1 benchmark
This dataset contains product reviews and metadata from Amazon, including 142.8 million reviews spanning May 1996 - July 2014.
41 papers · 5 benchmarks
V2V4Real is a large-scale real-world multi-modal dataset for V2V perception.
35 papers · 0 benchmarks
LM (LINEMOD)
The LM (Linemod) dataset is a valuable resource introduced by Stefan Hinterstoisser and colleagues in their research on model-based training, detection, and pose estimation of texture-less 3D objects in heavily cluttered scenes¹.
34 papers · 5 benchmarks
We introduce ACDC, the Adverse Conditions Dataset with Correspondences for training and testing semantic segmentation methods on adverse visual conditions.
31 papers · 5 benchmarks
AVD (Active Vision Dataset)
AVD focuses on simulating robotic vision tasks in everyday indoor environments using real imagery.
29 papers · 1 benchmark
Comic2k is a dataset used for cross-domain object detection which contains 2k comic images with image and instance-level annotations.
29 papers · 4 benchmarks
TechQA (The TechQA Dataset)
24 papers · 0 benchmarks
KdConv (Knowledge-driven Conversation)
KdConv is a Chinese multi-domain Knowledge-driven Conversation dataset, grounding the topics in multi-turn conversations to knowledge graphs.
22 papers · 0 benchmarks
Animal-Pose Dataset is an animal pose dataset to facilitate training and evaluation.
21 papers · 1 benchmark
VIDIT (Virtual Image Dataset for Illumination Transfer)
VIDIT is a reference evaluation benchmark and to push forward the development of illumination manipulation methods.
20 papers · 1 benchmark
Many existing datasets for lidar place recognition are solely representative of structured urban environments, and have recently been saturated in performance by deep learning based approaches.
18 papers · 1 benchmark
JESC (Japanese-English Subtitle Corpus)
Japanese-English Subtitle Corpus is a large Japanese-English parallel corpus covering the underrepresented domain of conversational dialogue.
17 papers · 0 benchmarks
VehicleX is a large-scale synthetic dataset.
16 papers · 0 benchmarks
CASIA V2 is a dataset for forgery classification.
15 papers · 0 benchmarks
CMU DoG (CMU Document Grounded Conversations Dataset)
This is a document grounded dataset for text conversations.
15 papers · 0 benchmarks
CrossNER is a cross-domain NER (Named Entity Recognition) dataset, a fully-labeled collection of NER data spanning over five diverse domains (Politics, Natural Science, Music, Literature, and Artificial Intelligence) with specialized…
15 papers · 1 benchmark
REFUGE Challenge (Retinal Fundus Glaucoma Challenge)
REFUGE Challenge provides a data set of 1200 fundus images with ground truth segmentations and clinical glaucoma labels, currently the largest existing one.
14 papers · 4 benchmarks
Multilingual Reuters (Multilingual Reuters Collection)
The Multilingual Reuters Collection dataset comprises over 11,000 articles from six classes in five languages, i.e., English (E), French (F), German (G), Italian (I), and Spanish (S).
13 papers · 0 benchmarks
Perspectrum is a dataset of claims, perspectives and evidence, making use of online debate websites to create the initial data collection, and augmenting it using search engines in order to expand and diversify the dataset.
11 papers · 1 benchmark
So2Sat LCZ42 consists of local climate zone (LCZ) labels of about half a million Sentinel-1 and Sentinel-2 image patches in 42 urban agglomerations (plus 10 additional smaller areas) across the globe.
11 papers · 1 benchmark
The CropAndWeed dataset is focused on the fine-grained identification of 74 relevant crop and weed species with a strong emphasis on data variability.
10 papers · 0 benchmarks
Office-Caltech-10 a standard benchmark for domain adaptation, which consists of Office 10 and Caltech 10 datasets.
10 papers · 1 benchmark

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.