{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-deep-learning-based-multimodal-fusion-model","title":"A deep learning based multimodal fusion model for skin lesion diagnosis using smartphone collected clinical images and metadata","arxiv_id":null,"date":"2022-10-04","proceeding":"Frontiers in Surgery 2022 10","authors":["Chubin Ou","Sitong Zhou","Ronghua Yang","Weili Jiang","Haoyang He","Wenjun Gan","Wentao Chen","Xinchi Qin","Wei Luo","Xiaobing Pi and Jiehua Li"],"abstract":"Introduction: Skin cancer is one of the most common types of cancer. An accessible tool to the public can help screening for malign lesion. We aimed to develop a deep learning model to classify skin lesion using clinical images and meta information collected from smartphones.\r\n\r\nMethods: A deep neural network was developed with two encoders for extracting information from image data and metadata. A multimodal fusion module with intra-modality self-attention and inter-modality cross-attention was proposed to effectively combine image features and meta features. The model was trained on tested on a public dataset and compared with other state-of-the-art methods using five-fold cross-validation.\r\n\r\nResults: Including metadata is shown to significantly improve a model's performance. Our model outperformed other metadata fusion methods in terms of accuracy, balanced accuracy and area under the receiver-operating characteristic curve, with an averaged value of 0.768±0.022, 0.775±0.022 and 0.947±0.007.\r\n\r\nConclusion: A deep learning model using smartphone collected images and metadata for skin lesion diagnosis was successfully developed. The proposed model showed promising performance and could be a potential tool for skin cancer screening.","url_abs":"https://www.frontiersin.org/journals/surgery/articles/10.3389/fsurg.2022.1029991/full","url_pdf":"https://www.frontiersin.org/journals/surgery/articles/10.3389/fsurg.2022.1029991/full","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"skin-lesion-classification","task_name":"Skin Lesion Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/skin-lesion-classification-on-pad-ufes-20","task":"Skin Lesion Classification","dataset":"PAD-UFES-20","model":"ViT+ Multimodality Cross-Attention Module","rank_in_archive_order":3,"of":3,"metrics":{"Balanced Accuracy":"0.775"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}