Home › Datasets › task › Domain Adaptation
Domain Adaptation datasets
archive 2025-07-28
96 datasets carry the task tag "Domain Adaptation" (the task itself: Domain Adaptation), ordered by the archive's paper count. Page 2 of 2: 48 shown of 96. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Domain Adaptation datasets 49–96 of 96
Adaptiope is a domain adaptation dataset with 123 classes in the three domains synthetic, product and real life.
9 papers · 0 benchmarks
Mindboggle is a large publicly available dataset of manually labeled brain MRI.
9 papers · 0 benchmarks
Under a close collaboration with an expert radiologist team of the Hospital Universitario San Cecilio, the COVIDGR-1.0 dataset of patients' anonymized X-ray images has been built.
8 papers · 2 benchmarks
The Cross-dataset Testbed is a Decaf7 based cross-dataset image classification dataset, which contains 40 categories of images from 3 domains: 3,847 images in Caltech256, 4,000 images in ImageNet, and 2,626 images for SUN.
7 papers · 0 benchmarks
Modern Office-31 is a refurbished version of the commonly used Office-31 dataset.
6 papers · 0 benchmarks
OOD-CV (Out Of Distribution Generalization in Computer Vision)
Enhancing the robustness of vision algorithms in real-world scenarios is challenging.
6 papers · 1 benchmark
PLABA (Plain Language Adaptation of Biomedical Abstracts)
Plain Language Adaptation of Biomedical Abstracts (PLABA) is a dataset designed for automatic adaptation that is both document- and sentence-aligned.
6 papers · 0 benchmarks
A dataset for evaluating text classification, domain adaptation, and active learning models.
5 papers · 0 benchmarks
Libri-Adapt aims to support unsupervised domain adaptation research on speech recognition models.
5 papers · 0 benchmarks
NLPeer is a multidomain corpus of more than 5k papers and 11k review reports from five different venues.
5 papers · 0 benchmarks
This dataset contains 114 individuals including 1824 images captured from two disjoint camera views.
5 papers · 1 benchmark
RoCoG-v2 (Robot Control Gestures) is a dataset intended to support the study of synthetic-to-real and ground-to-air video domain adaptation.
5 papers · 1 benchmark
The Sims4Action Dataset: a videogame-based dataset for Synthetic→Real domain adaptation for human activity recognition.
5 papers · 0 benchmarks
Unsupervised Domain Adaptation demonstrates great potential to mitigate domain shifts by transferring models from labeled source domains to unlabeled target domains.
4 papers · 3 benchmarks
CocoDoom is a collection of pre-recorded data extracted from Doom gaming sessions along with annotations in the MS Coco format.
4 papers · 0 benchmarks
The French National Institute of Geographical and Forest Information (IGN) has the mission to document and measure land-cover on French territory and provides referential geographical datasets, including high-resolution aerial images and…
4 papers · 1 benchmark
Human-Animal-Cartoon (HAC) dataset consists of seven actions (‘sleeping’, ‘watching tv’, ‘eating’, ‘drinking’, ‘swimming’, ‘running’, and ‘opening door’) performed by humans, animals, and cartoon figures, forming three different domains.
4 papers · 0 benchmarks
NHA12D (A New Pavement Crack Dataset)
NHA12D is an annotated pavement crack dataset that contains images with different viewpoints and pavements types.
4 papers · 0 benchmarks
We design an all-day semantic segmentation benchmark all-day CityScapes.
3 papers · 1 benchmark
The Five-Billion-Pixels dataset contains more than 5 billion labeled pixels of 150 high-resolution Gaofen-2 (4 m) satellite images, annotated in a 24-category system covering artificial-constructed, agricultural, and natural classes.
3 papers · 0 benchmarks
Long-term visual localization provides a benchmark datasets aimed at evaluating 6 DoF pose estimation accuracy over large appearance variations caused by changes in seasonal (summer, winter, spring, etc.) and illumination (dawn, day,…
3 papers · 0 benchmarks
Stanceosaurus is a corpus of 28,033 tweets in English, Hindi, and Arabic annotated with stance towards 251 misinformation claims.
3 papers · 0 benchmarks
Two datasets (synthetic and natural/real) containing simultaneously recorded egocentric and exocentric videos.
3 papers · 0 benchmarks
Youtbean is a dataset created from closed captions of YouTube product review videos.
3 papers · 0 benchmarks
A benchmark dataset for training and evaluating global cloud classification models.
2 papers · 0 benchmarks
The DAPlankton dataset consists of over 110k expert-labeled plankton images.
2 papers · 0 benchmarks
This is a dataset used to test deep learning-supported deep learning for fault diagnosis: - A digital twin model for a robot.
2 papers · 1 benchmark
The LeukemiaAttri dataset is a large-scale, multi-domain collection of microscopy images derived from leukemia patient samples, enriched with detailed morphological information.
2 papers · 2 benchmarks
MSDA (Multi-source domain adaptation dataset for text recognition)
5 domains: synthetic domain, document domain, street view domain, handwritten domain, and car license domain over five million images
2 papers · 2 benchmarks
Mila Simulated Floods Dataset is a 1.5 square km virtual world using the Unity3D game engine including urban, suburban and rural areas.
2 papers · 1 benchmark
Pano3D is a new benchmark for depth estimation from spherical panoramas.
2 papers · 0 benchmarks
Our proposed Synthetic-to-Real benchmark for more practical visual DA (termed S2RDA) includes two challenging transfer tasks of S2RDA-49 and S2RDA-MS-39.
2 papers · 0 benchmarks
XL-R2R (Cross-lingual Room-to-Room)
The XL-R2R dataset is built upon the R2R dataset and extends it with Chinese instructions.
2 papers · 0 benchmarks
The Apron Dataset focuses on training and evaluating classification and detection models for airport-apron logistics.
1 paper · 0 benchmarks
The goal of this project is to present two new datasets that seek to expand the capability of the Learning to See in the Dark Low-light enhancement CNN for the Canon 6D DSLR, and explore how the network performs when modified in various…
1 paper · 2 benchmarks
DRIFT (Domain-Adaptive Regression for Forest Monitoring)
The DRIFT dataset includes 25k image patches collected in five European countries sourced from aerial and nanosatellite image archives.
1 paper · 0 benchmarks
FGraDA (Fine-Grained Domain Adaptation Dataset)
Previous research for adapting a general neural machine translation (NMT) model into a specific domain usually neglects the diversity in translation within the same domain, which is a core problem for domain adaptation in real- world…
1 paper · 0 benchmarks
InfraParis is a novel and versatile dataset supporting multiple tasks across three modalities: RGB, depth, and infrared.
1 paper · 0 benchmarks
LiDAR-CS is a dataset for 3D object detection in real traffic.
1 paper · 0 benchmarks
Dataset release for the BMVC 2021 Paper "Few-Shot Domain Adaptation for Low Light RAW Image Enhancement" Abstract: Enhancing practical low light raw images is a difficult task due to severe noise and color distortions from short exposure…
1 paper · 2 benchmarks
Open MIC (Open Museum Identification Challenge)
Open MIC (Open Museum Identification Challenge) contains photos of exhibits captured in 10 distinct exhibition spaces of several museums which showcase paintings, timepieces, sculptures, glassware, relics, science exhibits, natural history…
1 paper · 0 benchmarks
OpenGDA is a benchmark for evaluating graph domain adaptation models.
1 paper · 0 benchmarks
We create Rwanda built-up regions dataset, a different and versatile in nature from previously available datasets.
1 paper · 0 benchmarks
SemanticUSL is a dataset for domain adaptation for LiDAR point cloud semantic segmentation.
1 paper · 0 benchmarks
This dataset is a large-scale synthetic dataset to simulate the attack scenario for a keystroke inference attack.
1 paper · 0 benchmarks
The Medical Translation Task of WMT 2014 addresses the problem of domain-specific and genre-specific machine translation.
1 paper · 0 benchmarks
Microarray gene expression data on 57 bladder samples from 5 batches.
1 paper · 0 benchmarks
fNIRS2MW (The Tufts fNIRS to Mental Workload Dataset)
The Tufts fNIRS to Mental Workload (fNIRS2MW) open-access dataset is a new dataset for building machine learning classifiers that can consume a short window (30 seconds) of multivariate fNIRS recordings and predict the mental workload…
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.