Browse State-of-the-Art › Multimodal Deep Learning
Multimodal Deep Learning
97 papers with code · 1 benchmark · 22 datasets archive 2025-07-28
Multimodal deep learning is a type of deep learning that combines information from multiple modalities, such as text, image, audio, and video, to make more accurate and comprehensive predictions. It involves training deep neural networks on data that includes multiple types of information and using the network to make predictions based on this combined data.
One of the key challenges in multimodal deep learning is how to effectively combine information from multiple modalities. This can be done using a variety of techniques, such as fusing the features extracted from each modality, or using attention mechanisms to weight the contribution of each modality based on its importance for the task at hand.
Multimodal deep learning has many applications, including image captioning, speech recognition, natural language processing, and autonomous vehicles. By combining information from multiple modalities, multimodal deep learning can improve the accuracy and robustness of models, enabling them to perform better in real-world scenarios where multiple types of information are present.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| CUB-200-2011 (1 row) | Two Branch Network (Text - Bert + Image - Nts-Net) | Are These Birds Similar: Learning Branched Networks for... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
22 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 97 papers with code (213 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
28 Mar 2023 7 repositories listedWe present LLaMA-Adapter, a lightweight adaption method to efficiently fine-tune LLaMA into an instruction-following model.
-
3 Oct 2023 6 repositories listed Syntology ran 7 of 14 samples · 7 unverifiedWe thus propose VIDAL-10M with Video, Infrared, Depth, Audio and their corresponding Language, naming as VIDAL-10M.
-
9 May 2023 3 repositories listed Syntology ran 24 of 34 samples · 10 unverified · 32 pointer-only (licence)We show that all combinations of paired data are not necessary to train such a joint embedding, and only image-paired data is sufficient to bind the modalities together.
-
17 Nov 2022 3 repositories listed Syntology ran 0 of 5 samples · 5 unverifiedThe dataset consists of 176 scenes with synchronized and calibrated LiDAR, camera, and radar sensors covering a 360-degree field of view.
-
23 Apr 2021 3 repositories listedThe proposed architecture utilizes an attention mechanism before fusing motion features and features representing the (static) visual content, i.
-
16 Jan 2020 3 repositories listedIn recent years, natural language descriptions are used to obtain information on discriminative parts of the object.
-
15 Jul 2019 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Classification of document images is a critical step for archival of old manuscripts, online subscription and administrative procedures.
-
14 Apr 2017 3 repositories listedWe introduce a novel framework for evaluating multimodal deep learning models with respect to their language understanding and generalization abilities.
-
16 Apr 2024 2 repositories listedTherefore, we introduce the first large-scale dataset in Vietnamese specializing in the ability to understand scene text, we call it ViTextVQA (\textbf{Vi}etnamese \textbf{Text}-based \textbf{V}isual \textbf{Q}uestion…
-
15 Nov 2023 2 repositories listed Syntology ran 6 of 11 samples · 5 unverifiedTechnological advances in medical data collection, such as high-throughput genomic sequencing and digital high-resolution histopathology, have contributed to the rising requirement for multimodal biomedical modelling,…
-
11 Nov 2023 2 repositories listedThe versatility of multimodal deep learning holds tremendous promise for advancing scientific research and practical applications.
-
2 Oct 2023 2 repositories listedOur MMDL system uses RETFound, a foundation model pre-trained on 1.
-
28 Jul 2021 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedMultimodal sentiment analysis aims to extract and integrate semantic information collected from multiple modalities to recognize the expressed emotions and sentiment in multimodal data.
-
18 May 2021 2 repositories listedWe propose a novel, efficient, modular and scalable framework for content based visual media retrieval systems by leveraging the power of Deep Learning which is flexible to work both for images and videos conjointly and…
-
24 Jul 2015 2 repositories listedRobust object recognition is a crucial ingredient of many, if not all, real-world robotics applications.
-
8 Apr 2025 1 repository listedMotivated by this, we introduce Gaze-CIFAR-10, a human gaze time-series dataset, along with a dual-sequence gaze encoder that models the precise sequential localization of human attention on distinct local attributes.
-
4 Mar 2025 1 repository listedMolecular subtyping of breast cancer is crucial for personalized treatment and prognosis.
-
9 Feb 2025 1 repository listedPDE foundation models utilize neural networks to train approximations to multiple differential equations simultaneously and are thus a general purpose solver that can be adapted to downstream tasks.
-
16 Jan 2025 1 repository listedOur analysis revealed that the MobileNet model achieved the highest accuracy of 99.
-
17 Dec 2024 1 repository listedThis unified lightweight model bridges the gap between various modalities and languages, enhancing its effectiveness in handling and retrieving multilingual and multimodal data.
-
27 Nov 2024 1 repository listedThis study introduces a multimodal deep-learning model leveraging mammogram datasets to evaluate breast cancer prediction.
-
22 Nov 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedConclusions: This study provides first evidence for the feasibility of using ECG data alongside clinical routine data for the real-time estimation and monitoring of laboratory value abnormalities, which could provide a…
-
23 Oct 2024 1 repository listedGlioblastoma is a highly aggressive form of brain cancer characterized by rapid progression and poor prognosis.
-
13 Sep 2024 1 repository listedIn this paper, we introduce PHemoNet, a fully hypercomplex network for multimodal emotion recognition from physiological signals.
-
6 Sep 2024 1 repository listedHowever, there is a gap between visual representation learning and textual semantic learning, and how to properly utilize the representation of two different modalities for clustering is still a big challenge.
-
30 Aug 2024 1 repository listedAdvancements in sixth-generation (6G) networks, coupled with the evolution of multimodal sensing in vehicle-to-everything (V2X) networks, have opened avenues for transformative research into multimodal-based artificial…
-
16 Aug 2024 1 repository listedRemarkably, our FoF achieves superior performance using only histopathology slides compared to existing multimodal methods.
-
20 Jun 2024 1 repository listedThis study introduces a novel geometric-based multimodal deep learning model for spatially embedded network representation learning, namely the regional spatial graph convolutional network (RSGCN).
-
14 Jun 2024 1 repository listedIt extends the well-known CIFAR 10/100 dataset with audio samples extracted from three audio corpora, and text data generated using the Gemma-7B Large Language Model (LLM).
-
3 Jun 2024 1 repository listedIn this paper, we introduce a pioneering multimodal DL-based approach for plant classification with automatic modality fusion.
Syntology lines on 7 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections