{"url":"/task/multimodal-deep-learning","name":"Multimodal Deep Learning","slug":"multimodal-deep-learning","description_markdown":"**Multimodal deep learning** is a type of deep learning that combines information from multiple modalities, such as text, image, audio, and video, to make more accurate and comprehensive predictions. It involves training deep neural networks on data that includes multiple types of information and using the network to make predictions based on this combined data.\r\n\r\nOne of the key challenges in multimodal deep learning is how to effectively combine information from multiple modalities. This can be done using a variety of techniques, such as fusing the features extracted from each modality, or using attention mechanisms to weight the contribution of each modality based on its importance for the task at hand.\r\n\r\nMultimodal deep learning has many applications, including image captioning, speech recognition, natural language processing, and autonomous vehicles. By combining information from multiple modalities, multimodal deep learning can improve the accuracy and robustness of models, enabling them to perform better in real-world scenarios where multiple types of information are present.","categories":[{"name":"Methodology","url":"/area/methodology"},{"name":"Natural Language Processing","url":"/area/natural-language-processing"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":213,"papers_with_code":97,"benchmarks":1,"benchmark_tables_in_archive":1,"benchmark_tables_shown":1,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":22,"subtasks":1,"parent_tasks":0},"benchmarks":[{"leaderboard":"/sota/multimodal-deep-learning-on-cub-200-2011","slug":"multimodal-deep-learning-on-cub-200-2011","dataset":"CUB-200-2011","dataset_url":"/dataset/cub-200-2011","rows_in_archive":1,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"Two Branch Network (Text - Bert + Image - Nts-Net)","paper_title":"Are These Birds Similar: Learning Branched Networks for Fine-grained Representations","paper_url":"/paper/are-these-birds-similar-learning-branched","paper_date":"2020-01-16","arxiv_id":null,"code_links":[{"title":"nicolalandro/ntsnet-cub200","url":"https://github.com/nicolalandro/ntsnet-cub200"},{"title":"Mind23-2/MindCode-101","url":"https://github.com/Mind23-2/MindCode-101/tree/main/ntsnet"},{"title":"artelabsuper/ivcnz-2019-bird","url":"https://gitlab.com/artelabsuper/ivcnz-2019-bird"}],"syntology":null}}],"datasets":[{"url":"/dataset/cub-200-2011","name":"CUB-200-2011","full_name":"Caltech-UCSD Birds-200-2011","num_papers_in_archive":2235},{"url":"/dataset/scienceqa","name":"ScienceQA","full_name":"Science Question Answering","num_papers_in_archive":339},{"url":"/dataset/quilt-1m","name":"QUILT-1M","full_name":"","num_papers_in_archive":17},{"url":"/dataset/olives-dataset","name":"OLIVES Dataset","full_name":"Ophthalmic Labels for Investigating Visual Eye Semantics","num_papers_in_archive":10},{"url":"/dataset/marine-video-kit","name":"MVK","full_name":"Marine Video Kit","num_papers_in_archive":9},{"url":"/dataset/bioscan-5m","name":"BIOSCAN-5M","full_name":"","num_papers_in_archive":4},{"url":"/dataset/mute","name":"MUTE","full_name":"Multimodal Bengali Hateful Memes Dataset","num_papers_in_archive":3},{"url":"/dataset/luma","name":"LUMA","full_name":"Learning from Uncertain and Multimodal Data","num_papers_in_archive":2},{"url":"/dataset/au-dataset-for-visuo-haptic-object","name":"AU Dataset for Visuo-Haptic Object Recognition for Robots","full_name":"","num_papers_in_archive":1},{"url":"/dataset/boombox","name":"Boombox","full_name":"","num_papers_in_archive":1},{"url":"/dataset/correlated-corrupted-dataset","name":"Correlated Corrupted Dataset","full_name":"CCD","num_papers_in_archive":1},{"url":"/dataset/gaze-cifar-10","name":"Gaze-CIFAR-10","full_name":"","num_papers_in_archive":1},{"url":"/dataset/gebid","name":"GeBiD","full_name":"Geometric shapes Bimodal Dataset","num_papers_in_archive":1},{"url":"/dataset/mminstruct-gpt4v","name":"MMInstruct-GPT4V","full_name":"MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity","num_papers_in_archive":1},{"url":"/dataset/multimodal-pisa","name":"Multimodal PISA","full_name":"Multimodal Piano Skills Assessment","num_papers_in_archive":1},{"url":"/dataset/rebus","name":"REBUS","full_name":"A Robust Evaluation Benchmark of Understanding Symbols","num_papers_in_archive":1},{"url":"/dataset/ucd","name":"Uncorrelated Corrupted Dataset","full_name":"UCD","num_papers_in_archive":1},{"url":"/dataset/webli","name":"WebLI","full_name":"Web Language Image","num_papers_in_archive":1},{"url":"/dataset/mimic-meme-dataset","name":"MIMIC Meme Dataset","full_name":"Misogyny Identification in Multimodal Internet Content in Hindi-English Code-Mix Language","num_papers_in_archive":0},{"url":"/dataset/mudestreda","name":"Mudestreda","full_name":"Mudestreda Multimodal Device State Recognition Dataset","num_papers_in_archive":0},{"url":"/dataset/sf-tl54","name":"SF-TL54: A Thermal Facial Landmark Dataset with Visual Pairs","full_name":"SF-TL54: A Thermal Facial Landmark Dataset with Visual Pairs","num_papers_in_archive":0},{"url":"/dataset/vidimu-multimodal-video-and-imu-kinematic","name":"VIDIMU: Multimodal video and IMU kinematic dataset on daily life activities using affordable devices","full_name":"https://zenodo.org/record/8210563","num_papers_in_archive":0}],"subtasks":[{"url":"/task/multimodal-text-and-image-classification","name":"Multimodal Text and Image Classification"}],"parent_tasks":[],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":30,"of":97,"tagged_in_all":213,"items":[{"url":"/paper/llama-adapter-efficient-fine-tuning-of","title":"LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention","date":"2023-03-28","arxiv_id":"2303.16199","repositories_listed":7,"syntology":null},{"url":"/paper/languagebind-extending-video-language","title":"LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment","date":"2023-10-03","arxiv_id":"2310.01852","repositories_listed":6,"syntology":{"n":14,"n_ran":7,"n_unverified":7,"n_pointer_only":0}},{"url":"/paper/imagebind-one-embedding-space-to-bind-them","title":"ImageBind: One Embedding Space To Bind Them All","date":"2023-05-09","arxiv_id":"2305.05665","repositories_listed":3,"syntology":{"n":34,"n_ran":24,"n_unverified":10,"n_pointer_only":32}},{"url":"/paper/aimotive-dataset-a-multimodal-dataset-for","title":"aiMotive Dataset: A Multimodal Dataset for Robust Autonomous Driving with Long-Range Perception","date":"2022-11-17","arxiv_id":"2211.09445","repositories_listed":3,"syntology":{"n":5,"n_ran":0,"n_unverified":5,"n_pointer_only":0}},{"url":"/paper/supervised-video-summarization-via-multiple","title":"Supervised Video Summarization via Multiple Feature Sets with Parallel Attention","date":"2021-04-23","arxiv_id":"2104.11530","repositories_listed":3,"syntology":null},{"url":"/paper/are-these-birds-similar-learning-branched","title":"Are These Birds Similar: Learning Branched Networks for Fine-grained Representations","date":"2020-01-16","arxiv_id":null,"repositories_listed":3,"syntology":null},{"url":"/paper/multimodal-deep-networks-for-text-and-image","title":"Multimodal deep networks for text and image-based document classification","date":"2019-07-15","arxiv_id":"1907.06370","repositories_listed":3,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":1}},{"url":"/paper/shapeworld-a-new-test-methodology-for","title":"ShapeWorld - A new test methodology for multimodal language understanding","date":"2017-04-14","arxiv_id":"1704.04517","repositories_listed":3,"syntology":null},{"url":"/paper/vitextvqa-a-large-scale-visual-question","title":"ViTextVQA: A Large-Scale Visual Question Answering Dataset for Evaluating Vietnamese Text Comprehension in Images","date":"2024-04-16","arxiv_id":"2404.10652","repositories_listed":2,"syntology":null},{"url":"/paper/healnet-hybrid-multi-modal-fusion-for","title":"HEALNet: Multimodal Fusion for Heterogeneous Biomedical Data","date":"2023-11-15","arxiv_id":"2311.09115","repositories_listed":2,"syntology":{"n":11,"n_ran":6,"n_unverified":5,"n_pointer_only":0}},{"url":"/paper/asymmetric-contrastive-multimodal-learning","title":"Advancing Drug Discovery with Enhanced Chemical Understanding via Asymmetric Contrastive Multimodal Learning","date":"2023-11-11","arxiv_id":"2311.06456","repositories_listed":2,"syntology":null},{"url":"/paper/hymnet-a-multimodal-deep-learning-system-for","title":"HyMNet: a Multimodal Deep Learning System for Hypertension Classification using Fundus Photographs and Cardiometabolic Risk Factors","date":"2023-10-02","arxiv_id":"2310.01099","repositories_listed":2,"syntology":null},{"url":"/paper/bi-bimodal-modality-fusion-for-correlation","title":"Bi-Bimodal Modality Fusion for Correlation-Controlled Multimodal Sentiment Analysis","date":"2021-07-28","arxiv_id":"2107.13669","repositories_listed":2,"syntology":{"n":2,"n_ran":2,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/a-multimodal-deep-learning-framework-for","title":"A multimodal deep learning framework for scalable content based visual media retrieval","date":"2021-05-18","arxiv_id":"2105.08665","repositories_listed":2,"syntology":null},{"url":"/paper/multimodal-deep-learning-for-robust-rgb-d","title":"Multimodal Deep Learning for Robust RGB-D Object Recognition","date":"2015-07-24","arxiv_id":"1507.06821","repositories_listed":2,"syntology":null},{"url":"/paper/gaze-guided-learning-avoiding-shortcut-bias","title":"Gaze-Guided Learning: Avoiding Shortcut Bias in Visual Classification","date":"2025-04-08","arxiv_id":"2504.05583","repositories_listed":1,"syntology":null},{"url":"/paper/multimodal-deep-learning-for-subtype","title":"Multimodal Deep Learning for Subtype Classification in Breast Cancer Using Histopathological Images and Gene Expression Data","date":"2025-03-04","arxiv_id":"2503.02849","repositories_listed":1,"syntology":null},{"url":"/paper/a-multimodal-pde-foundation-model-for","title":"A Multimodal PDE Foundation Model for Prediction and Scientific Text Descriptions","date":"2025-02-09","arxiv_id":"2502.06026","repositories_listed":1,"syntology":null},{"url":"/paper/multimodal-marvels-of-deep-learning-in","title":"Multimodal Marvels of Deep Learning in Medical Diagnosis: A Comprehensive Review of COVID-19 Detection","date":"2025-01-16","arxiv_id":"2501.09506","repositories_listed":1,"syntology":null},{"url":"/paper/clasp-contrastive-language-speech-pretraining","title":"CLASP: Contrastive Language-Speech Pretraining for Multilingual Multimodal Information Retrieval","date":"2024-12-17","arxiv_id":"2412.13071","repositories_listed":1,"syntology":null},{"url":"/paper/frozen-large-scale-pretrained-vision-language","title":"Frozen Large-scale Pretrained Vision-Language Models are the Effective Foundational Backbone for Multimodal Breast Cancer Prediction","date":"2024-11-27","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/cardiolab-laboratory-values-estimation-and","title":"CardioLab: Laboratory Values Estimation and Monitoring from Electrocardiogram Signals -- A Multimodal Deep Learning Approach","date":"2024-11-22","arxiv_id":"2411.14886","repositories_listed":1,"syntology":{"n":1,"n_ran":1,"n_unverified":0,"n_pointer_only":0}},{"url":"/paper/uncovering-the-genetic-basis-of-glioblastoma","title":"Uncovering the Genetic Basis of Glioblastoma Heterogeneity through Multimodal Analysis of Whole Slide Images and RNA Sequencing Data","date":"2024-10-23","arxiv_id":"2410.18710","repositories_listed":1,"syntology":null},{"url":"/paper/phemonet-a-multimodal-network-for","title":"PHemoNet: A Multimodal Network for Physiological Signals","date":"2024-09-13","arxiv_id":"2410.00010","repositories_listed":1,"syntology":null},{"url":"/paper/dual-level-cross-modal-contrastive-clustering","title":"Dual-Level Cross-Modal Contrastive Clustering","date":"2024-09-06","arxiv_id":"2409.04561","repositories_listed":1,"syntology":null},{"url":"/paper/mvx-vit-multimodal-collaborative-perception-1","title":"MVX-ViT: Multimodal Collaborative Perception for 6G V2X Network Management Decisions Using Vision Transformer.","date":"2024-08-30","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/focus-on-focus-focus-oriented-representation","title":"Focus on Focus: Focus-oriented Representation Learning and Multi-view Cross-modal Alignment for Glioma Grading","date":"2024-08-16","arxiv_id":"2408.08527","repositories_listed":1,"syntology":null},{"url":"/paper/modeling-of-spatially-embedded-networks-via","title":"Modeling of spatially embedded networks via regional spatial graph convolutional networks","date":"2024-06-20","arxiv_id":null,"repositories_listed":1,"syntology":null},{"url":"/paper/luma-a-benchmark-dataset-for-learning-from","title":"LUMA: A Benchmark Dataset for Learning from Uncertain and Multimodal Data","date":"2024-06-14","arxiv_id":"2406.09864","repositories_listed":1,"syntology":null},{"url":"/paper/automatic-fused-multimodal-deep-learning-for","title":"Automatic Fused Multimodal Deep Learning for Plant Identification","date":"2024-06-03","arxiv_id":"2406.01455","repositories_listed":1,"syntology":null}],"syntology_records":7,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}