Home › Datasets › task › Multi-Task Learning

Multi-Task Learning datasets

archive 2025-07-28

59 datasets carry the task tag "Multi-Task Learning" (the task itself: Multi-Task Learning), ordered by the archive's paper count. Page 1 of 2: 48 shown of 59. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Multi-Task Learning datasets 1–48 of 59

Cityscapes is a large-scale database which focuses on semantic understanding of urban street scenes.
3,702 papers · 51 benchmarks
CelebA (CelebFaces Attributes Dataset)
CelebFaces Attributes dataset contains 202,599 face images of the size 178×218 from 10,177 celebrities, each annotated with 40 binary labels indicating facial attributes like hair color, gender and age.
3,477 papers · 17 benchmarks
NYUv2 (NYU-Depth V2)
The NYU-Depth V2 data set is comprised of video sequences from a variety of indoor scenes as recorded by both the RGB and Depth cameras from the Microsoft Kinect.
986 papers · 16 benchmarks
The UTKFace dataset is a large-scale face dataset with long age span (range from 0 to 116 years old).
243 papers · 4 benchmarks
ChestX-ray14 is a medical imaging dataset which comprises 112,120 frontal-view X-ray images of 30,805 (collected from the year of 1992 to 2015) unique patients with the text-mined fourteen common disease labels, mined from the text…
237 papers · 6 benchmarks
Clotho is an audio captioning dataset, consisting of 4981 audio samples, and each audio sample has five captions (a total of 24 905 captions).
202 papers · 3 benchmarks
A new benchmark that spans concepts in justice, well-being, duties, virtues, and commonsense morality.
183 papers · 1 benchmark
Aff-Wild2 is a large-scale in-the-wild database and an extension of the Aff-Wild dataset for affect recognition.
142 papers · 2 benchmarks
For many fundamental scene understanding tasks, it is difficult or impossible to obtain per-pixel ground truth labels from real images.
108 papers · 4 benchmarks
A dataset for robot navigation task and more.
107 papers · 0 benchmarks
KP20k is a large-scale scholarly articles dataset with 528K articles for training, 20K articles for validation and 20K articles for testing.
87 papers · 3 benchmarks
QM9 provides quantum chemical properties (at DFT level) for a relevant, consistent, and comprehensive chemical space of small organic molecules.
76 papers · 9 benchmarks
An open-source simulated benchmark for meta-reinforcement learning and multi-task learning consisting of 50 distinct robotic manipulation tasks.
73 papers · 0 benchmarks
CELEX database comprises three different searchable lexical databases, Dutch, English and German.
59 papers · 0 benchmarks
The Wireframe dataset consists of 5,462 images (5,000 for training, 462 for test) of indoor and outdoor man-made scenes.
57 papers · 2 benchmarks
KuaiRand is an unbiased sequential recommendation dataset collected from the recommendation logs of the video-sharing mobile app, Kuaishou (快手).
42 papers · 1 benchmark
THCHS-30 is a free Chinese speech database THCHS-30 that can be used to build a full-fledged Chinese speech recognition system.
34 papers · 0 benchmarks
A new multitask action quality assessment (AQA) dataset, the largest to date, comprising of more than 1600 diving samples; contains detailed annotations for fine-grained action recognition, commentary generation, and estimating the AQA…
33 papers · 2 benchmarks
CH-SIMS is a Chinese single- and multimodal sentiment analysis dataset which contains 2,281 refined video segments in the wild with both multimodal and independent unimodal annotations.
25 papers · 1 benchmark
WIDER (Web Image Dataset for Event Recognition)
WIDER is a dataset for complex event recognition from static images.
22 papers · 1 benchmark
CAL500 (Computer Audition Lab 500)
CAL500 (Computer Audition Lab 500) is a dataset aimed for evaluation of music information retrieval systems.
21 papers · 0 benchmarks
Consists of 1106 action samples from seven actions with quality scores as measured by expert human judges.
20 papers · 1 benchmark
A suite of open-source federated datasets, a rigorous evaluation framework, and a set of reference implementations, all geared towards capturing the obstacles and intricacies of practical federated environments.
19 papers · 0 benchmarks
JuICe (JuICe Dataset)
JuICe is a corpus of 1.5 million examples with a curated test set of 3.7K instances based on online programming assignments.
16 papers · 0 benchmarks
SemArt is a multi-modal dataset for semantic art understanding.
16 papers · 0 benchmarks
GCDC (Grammarly Corpus of Discourse Coherence)
A corpus of real-world texts.
13 papers · 2 benchmarks
FSDKaggle2018 is an audio dataset containing 11,073 audio files annotated with 41 labels of the AudioSet Ontology.
12 papers · 1 benchmark
The CropAndWeed dataset is focused on the fine-grained identification of 74 relevant crop and weed species with a strong emphasis on data variability.
10 papers · 0 benchmarks
SkillSpan (Hard and Soft Skill Extraction from English Job Postings)
SkillSpan is a dataset for Skill Extraction (SE).
10 papers · 0 benchmarks
NCLS (Neural Cross-Lingual Summarization Corpora)
Presents two high-quality large-scale CLS datasets based on existing monolingual summarization datasets.
9 papers · 0 benchmarks
Presents half a million samples and structured meta-data to encourage further research and societal engagement.
9 papers · 1 benchmark
The Video-based Multimodal Summarization with Multimodal Output (VMSMO) corpus consists of 184,920 document-summary pairs, with 180,000 training pairs, 2,460 validation and test pairs.
8 papers · 0 benchmarks
FSDKaggle2019 is an audio dataset containing 29,266 audio files annotated with 80 labels of the AudioSet Ontology.
5 papers · 0 benchmarks
PGDP5K (Plane Geometry Diagram Parsing Dataset)
PGDP5K is a dataset consisting of 5000 diagram samples composed of 16 shapes, covering 5 positional relations, 22 symbol types and 6 text types, labeled with more fine-grained annotations at primitive level, including primitive classes,…
5 papers · 1 benchmark
WWW Crowd provides 10,000 videos with over 8 million frames from 8,257 diverse scenes, therefore offering a comprehensive dataset for the area of crowd understanding.
5 papers · 0 benchmarks
CS (Chinese Simile)
This dataset is constructed and based on the online free-access fictions that are tagged with sci-fi, urban novel, love story, youth, etc.
3 papers · 0 benchmarks
This is a dataset for segmentation and classification of epistemic activities in diagnostic reasoning texts.
3 papers · 0 benchmarks
HSD (Honda Scenes Dataset)
An annotated dataset is released to enable dynamic scene classification that includes 80 hours of diverse high quality driving video data clips collected in the San Francisco Bay area.
3 papers · 0 benchmarks
Publicly available dataset in the hotel domain (50M versus 0.9M) and additionally, the largest recommendation dataset in a single domain and with textual reviews (50M versus 22M).
3 papers · 0 benchmarks
MMDB (Multimodal Dyadic Behavior)
Multimodal Dyadic Behavior (MMDB) dataset is a unique collection of multimodal (video, audio, and physiological) recordings of the social and communicative behavior of toddlers.
3 papers · 0 benchmarks
OSAI introduces OpenTTGames - an open dataset aimed at evaluation of different computer vision tasks in Table Tennis: ball detection, semantic segmentation of humans, table and scoreboard and fast in-game events spotting.
3 papers · 0 benchmarks
RoboPianist is a benchmarking suite for high-dimensional control, targeted at testing high spatial and temporal precision, coordination, and planning, all with an underactuated system frequently making-and-breaking contacts.
3 papers · 0 benchmarks
Huggingface Datasets is a great library, but it lacks standardization, and datasets require preprocessing work to be used interchangeably.
3 papers · 0 benchmarks
CQR (Contextual Query Rewrite)
CQR is an extension to the Stanford Dialogue Corpus.
2 papers · 0 benchmarks
ExHVV is a novel dataset that offers natural language explanations of connotative roles for three types of entities -- heroes, villains, and victims, encompassing 4,680 entities present in 3K memes.
2 papers · 0 benchmarks
The Cifar10Mnist dataset is created using CIFAR-10 and MNIST data sources.
1 paper · 0 benchmarks
FIW-MM (Families In Wild Multimedia)
A large-scale dataset for recognizing kinship in multimedia which extend FIW with multimedia data (i.e., video, audio, and contextual transcripts).
1 paper · 0 benchmarks
A large-scale machine comprehension dataset (based on the COCO images and captions).
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.