Home › Datasets › task › Data Augmentation
Data Augmentation datasets
archive 2025-07-28
63 datasets carry the task tag "Data Augmentation" (the task itself: Data Augmentation), ordered by the archive's paper count. Page 1 of 2: 48 shown of 63. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Data Augmentation datasets 1–48 of 63
description withheld: archive row vandalised before snapshot
16,145 papers · 91 benchmarks
The ImageNet dataset contains 14,197,122 annotated images according to the WordNet hierarchy.
15,430 papers · 52 benchmarks
The ESC-50 dataset is a labeled collection of 2000 environmental audio recordings suitable for benchmarking methods of environmental sound classification.
387 papers · 4 benchmarks
ImageNet-Sketch data set consists of 50,889 images, approximately 50 images for each of the 1000 ImageNet classes.
268 papers · 3 benchmarks
MuST-C currently represents the largest publicly available multilingual corpus (one-to-many) for speech translation.
216 papers · 2 benchmarks
Clotho is an audio captioning dataset, consisting of 4981 audio samples, and each audio sample has five captions (a total of 24 905 captions).
202 papers · 3 benchmarks
The IAM database contains 13,353 images of handwritten lines of text created by 657 writers.
198 papers · 1 benchmark
MathQA significantly enhances the AQuA dataset with fully-specified operational programs.
159 papers · 1 benchmark
PAWS (Paraphrase Adversaries from Word Scrambling)
Paraphrase Adversaries from Word Scrambling (PAWS) is a dataset contains 108,463 human-labeled and 656k noisily labeled pairs that feature the importance of modeling structure, context, and word order information for the problem of…
159 papers · 0 benchmarks
Urban Sound 8K is an audio dataset that contains 8732 labeled sound excerpts (<=4s) of urban sounds from 10 classes: airconditioner, carhorn, childrenplaying, dogbark, drilling, engingeidling, gunshot, jackhammer, siren, and streetmusic.
147 papers · 1 benchmark
GAP is a gender-balanced dataset containing 8,908 coreference-labeled pairs of (ambiguous pronoun, antecedent name), sampled from Wikipedia and released by Google AI Language for the evaluation of coreference resolution in practical…
104 papers · 0 benchmarks
The PROMISE12 dataset was made available for the MICCAI 2012 prostate segmentation challenge.
84 papers · 2 benchmarks
SParC (Semantic Parsing in Context)
SParC is a large-scale dataset for complex, cross-domain, and context-dependent (multi-turn) semantic parsing and text-to-SQL task (interactive natural language interfaces for relational databases).
59 papers · 2 benchmarks
Europarl-ST is a multilingual Spoken Language Translation corpus containing paired audio-text samples for SLT from and into 9 European languages, for a total of 72 different translation directions.
57 papers · 0 benchmarks
The xBD dataset contains over 45,000KM2 of polygon labeled pre and post disaster imagery.
51 papers · 2 benchmarks
The Synthesized Lakh (Slakh) Dataset is a dataset for audio source separation that is synthesized from the Lakh MIDI Dataset v0.1 using professional-grade sample-based virtual instruments.
38 papers · 3 benchmarks
Contains hundreds of frontal view X-rays and is the largest public resource for COVID-19 image and prognostic data, making it a necessary resource to develop and evaluate tools to aid in the treatment of COVID-19.
35 papers · 1 benchmark
MeQSum is a dataset for medical question summarization.
33 papers · 1 benchmark
MedQuAD (Medical Question Answering Dataset)
MedQuAD includes 47,457 medical question-answer pairs created from 12 NIH websites (e.g.
32 papers · 0 benchmarks
The HELP dataset is an automatically created natural language inference (NLI) dataset that embodies the combination of lexical and logical inferences focusing on monotonicity (i.e., phrase replacement-based reasoning).
30 papers · 1 benchmark
The MRNet dataset consists of 1,370 knee MRI exams performed at Stanford University Medical Center.
29 papers · 1 benchmark
CONAN (COunter NArratives through Nichesourcing)
COunter NArratives through Nichesourcing (CONAN) is a dataset that consists of 4,078 pairs over the 3 languages.
27 papers · 0 benchmarks
Evidence Inference is a corpus for this task comprising 10,000+ prompts coupled with full-text articles describing RCTs.
27 papers · 0 benchmarks
CCPD (Chinese City Parking Dataset)
The Chinese City Parking Dataset (CCPD) is a dataset for license plate detection and recognition.
24 papers · 0 benchmarks
KdConv (Knowledge-driven Conversation)
KdConv is a Chinese multi-domain Knowledge-driven Conversation dataset, grounding the topics in multi-turn conversations to knowledge graphs.
22 papers · 0 benchmarks
ETHOS (multi-labEl haTe speecH detectiOn dataSet)
ETHOS is a hate speech detection dataset.
20 papers · 2 benchmarks
The George Washington dataset contains 20 pages of letters written by George Washington and his associates in 1755 and thereby categorized into historical collection.
20 papers · 0 benchmarks
InfoTabS comprises of human-written textual hypotheses based on premises that are tables extracted from Wikipedia info-boxes.
18 papers · 0 benchmarks
MED (Monotonicity Entailment Dataset)
MED is a new evaluation dataset that covers a wide range of monotonicity reasoning that was created by crowdsourcing and collected from linguistics publications.
18 papers · 1 benchmark
ORCAS is a click-based dataset.
18 papers · 0 benchmarks
RPC (Retail Product Checkout)
RPC is a large-scale retail product checkout dataset and collects 200 retail SKUs.
17 papers · 0 benchmarks
VehicleX is a large-scale synthetic dataset.
16 papers · 0 benchmarks
BRATS 2016 is a brain tumor segmentation dataset.
13 papers · 0 benchmarks
Logo-2K+:A Large-Scale Logo Dataset for Scalable Logo Classification The Logo-2K+ dataset contains a diverse range of logo classes from real-world logo images.
13 papers · 0 benchmarks
CRD3 (Critical Role Dungeons and Dragons Dataset)
The dataset is collected from 159 Critical Role episodes transcribed to text dialogues, consisting of 398,682 turns.
9 papers · 0 benchmarks
EgoHOS (Fine-Grained Egocentric Hand-Object Segmentation Dataset)
EgoHOS is a labeled dataset consisting of 11243 egocentric images with per-pixel segmentation labels of hands and objects being interacted with during a diverse array of daily activities.
9 papers · 0 benchmarks
Europarl-ASR (EN) is a 1300-hour English-language speech and text corpus of parliamentary debates for (streaming) Automatic Speech Recognition training and benchmarking, speech data filtering and speech data verbatimization, based on…
8 papers · 2 benchmarks
Contains 446,684 images annotated by humans that cover 43 incidents across a variety of scenes.
8 papers · 0 benchmarks
The Hotels-50K dataset consists of over 1 million images from 50,000 different hotels around the world.
7 papers · 0 benchmarks
Created from endoscopic video feeds of real-world surgical procedures.
5 papers · 0 benchmarks
word2word contains easy-to-use word translations for 3,564 language pairs.
5 papers · 0 benchmarks
Amazon Fine Foods is a dataset that consists of reviews of fine foods from amazon.
4 papers · 0 benchmarks
Bentham manuscripts refers to a large set of documents that were written by the renowned English philosopher and reformer Jeremy Bentham (1748-1832).
4 papers · 1 benchmark
The BirdVox-full-night dataset contains 6 audio recordings, each about ten hours in duration.
3 papers · 0 benchmarks
Konzil dataset was created by specialists of the University of Greifswald.
3 papers · 0 benchmarks
The LITIS-Rouen dataset is a dataset for audio scenes.
3 papers · 0 benchmarks
Patzig contains handwritten texts written in modern German.
3 papers · 0 benchmarks
Ricordi contains handwritten texts written in Italian.
3 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.