Home › Datasets › task › Multimodal Deep Learning
Multimodal Deep Learning datasets
archive 2025-07-28
22 datasets carry the task tag "Multimodal Deep Learning" (the task itself: Multimodal Deep Learning), ordered by the archive's paper count. Page 1 of 1: 22 shown of 22. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Multimodal Deep Learning datasets 1–22 of 22
The Caltech-UCSD Birds-200-2011 (CUB-200-2011) dataset is the most widely-used dataset for fine-grained visual categorization task.
2,235 papers · 47 benchmarks
Science Question Answering (ScienceQA) is a new benchmark that consists of 21,208 multimodal multiple choice questions with diverse science topics and annotations of their answers with corresponding lectures and explanations.
339 papers · 1 benchmark
Recent accelerations in multi-modal applications have been made possible with the plethora of image and text data available online.
17 papers · 0 benchmarks
OLIVES Dataset (Ophthalmic Labels for Investigating Visual Eye Semantics)
Clinical diagnosis of the eye is performed over multifarious data modalities including scalar clinical labels, vectorized biomarkers, two-dimensional fundus images, and three-dimensional Optical Coherence Tomography (OCT) scans.
10 papers · 0 benchmarks
The dataset contains single-shot videos taken from moving cameras in underwater environments.
9 papers · 1 benchmark
As part of an ongoing worldwide effort to comprehend and monitor insect biodiversity, we present the BIOSCAN-5M Insect dataset to the machine learning community.
4 papers · 0 benchmarks
MUTE (Multimodal Bengali Hateful Memes Dataset)
MUTE This is the first open-source Bengali Hateful Meme dataset, consisting of around 4200 memes annotated with two labels: hate and not hate.
3 papers · 0 benchmarks
LUMA (Learning from Uncertain and Multimodal Data)
LUMA is a multimodal dataset that consists of audio, image, and text modalities.
2 papers · 0 benchmarks
Multimodal object recognition is still an emerging field.
1 paper · 0 benchmarks
Boombox is a multi-modal dataset for visual reconstruction from acoustic vibrations.
1 paper · 0 benchmarks
Correlated Corrupted Dataset is an evaluation set that consists of realistic visible-infrared (V-I) corruptions allowing for models' corruption robustness evaluation.
1 paper · 0 benchmarks
We construct Gaze-CIFAR-10, a gaze-augmented image dataset based on the standard CIFAR-10 benchmark, enhanced with human eye-tracking annotations collected using the HTC VIVE Pro Eye headset.
1 paper · 1 benchmark
GeBiD (Geometric shapes Bimodal Dataset)
We provide a custom synthetic bimodal dataset, called GeBiD, designed specifically for the comparison of the joint- and cross-generative capabilities of Multimodal Variational Autoencoders.
1 paper · 0 benchmarks
MMInstruct-GPT4V (MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity)
Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations: 1.
1 paper · 0 benchmarks
Dataset for multimodal skills assessment focusing on assessing piano player’s skill level.
1 paper · 3 benchmarks
REBUS (A Robust Evaluation Benchmark of Understanding Symbols)
Recent advances in large language models have led to the development of multimodal LLMs (MLLMs), which take both image data and text as an input.
1 paper · 1 benchmark
Uncorrelated Corrupted Dataset is an evaluation set that consists of realistic visible-infrared (V-I) corruptions allowing for models' corruption robustness evaluation.
1 paper · 0 benchmarks
WebLI (Web Language Image)
WebLI (Web Language Image) is a web-scale multilingual image-text dataset, designed to support Google’s vision-language research, such as the large-scale pre-training for image understanding, image captioning, visual question answering,…
1 paper · 0 benchmarks
MIMIC Meme Dataset (Misogyny Identification in Multimodal Internet Content in Hindi-English Code-Mix Language)
This dataset endeavors to fill the research void by presenting a meticulously curated collection of misogynistic memes in a code-mixed language of Hindi and English.
0 papers · 0 benchmarks
Mudestreda (Mudestreda Multimodal Device State Recognition Dataset)
Mudestreda Multimodal Device State Recognition Dataset obtained from real industrial milling device with Time Series and Image Data for Classification, Regression, Anomaly Detection, Remaining Useful Life (RUL) estimation, Signal Drift…
0 papers · 0 benchmarks
Facial landmark detection is a cornerstone in many facial analysis tasks such as face recognition, drowsiness detection, and facial expression recognition.
0 papers · 0 benchmarks
Human activity recognition and clinical biomechanics are challenging problems in physical telerehabilitation medicine.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.