{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/arxiv-2508-03742","title":"Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-language Pre-training","arxiv_id":"2508.03742","date":"2025-08-01","proceeding":"ICCV","authors":["Weiwei Cao","Jianpeng Zhang","Zhongyi Shui","Sinuo Wang","Zeli Chen","Xi Li","Le Lu","Xianghua Ye","Tingbo Liang","Qi Zhang","Ling Zhang"],"abstract":"Vision-language pre-training (VLP) has great potential for developing multifunctional and general medical diagnostic capabilities. However, aligning medical images with a low signal-to-noise ratio (SNR) to reports with a high SNR presents a semantic density gap, leading to visual alignment bias. In this paper, we propose boosting vision semantic density to improve alignment effectiveness. On one hand, we enhance visual semantics through disease-level vision contrastive learning, which strengthens the model's ability to differentiate between normal and abnormal samples for each anatomical structure. On the other hand, we introduce an anatomical normality modeling method to model the distribution of normal samples for each anatomy, leveraging VQ-VAE for reconstructing normal vision embeddings in the latent space. This process amplifies abnormal signals by leveraging distribution shifts in abnormal samples, enhancing the model's perception and discrimination of abnormal attributes. The enhanced visual representation effectively captures the diagnostic-relevant semantics, facilitating more efficient and accurate alignment with the diagnostic report. We conduct extensive experiments on two chest CT datasets, CT-RATE and Rad-ChestCT, and an abdominal CT dataset, MedVL-CT69K, and comprehensively evaluate the diagnosis performance across multiple tasks in the chest and abdominal CT scenarios, achieving state-of-the-art zero-shot performance. Notably, our method achieved an average AUC of 84.9% across 54 diseases in 15 organs, significantly surpassing existing methods. Additionally, we demonstrate the superior transfer learning capabilities of our pre-trained model. Code is available at https://github.com/alibaba-damo-academy/ViSD-Boost.","url_abs":"https://arxiv.org/abs/2508.03742","url_pdf":"https://arxiv.org/pdf/2508.03742","source":{"archive":null,"snapshot":"2025-07-28","note":"not in the Papers with Code archive (frozen at the snapshot)","row_kind":"graph","title_abstract_authors_date":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)"},"code_links":[],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2508.03742","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2508.03742"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"mentioned_in_github":null,"is_official":null,"provenance":"deterministic:regex_extraction","mentioned_in_paper":null,"url":"https://github.com/alibaba-damo-academy/ViSD-Boost","reach":null}],"summary":{"ran":6,"unverified":1},"by_repo_kind":{"found_in_text":{"samples":7,"ran":6,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":7,"samples":[{"code_sha256_prefix":"ed3aafb236b56fa3","entry":"Attention","repo":"alibaba-damo-academy/ViSD-Boost","repo_kind":"found_in_text","path":"lavis/utils/model.py","file_url":"https://github.com/alibaba-damo-academy/ViSD-Boost/blob/HEAD/lavis/utils/model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"ed3aafb236b56fa3"}},{"code_sha256_prefix":"1c46be44162564f3","entry":"ExponentialMovingAverage","repo":"alibaba-damo-academy/ViSD-Boost","repo_kind":"found_in_text","path":"lavis/utils/model.py","file_url":"https://github.com/alibaba-damo-academy/ViSD-Boost/blob/HEAD/lavis/utils/model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"1c46be44162564f3"}},{"code_sha256_prefix":"41a29c1c6528fdf6","entry":"Transformer","repo":"alibaba-damo-academy/ViSD-Boost","repo_kind":"found_in_text","path":"lavis/utils/model.py","file_url":"https://github.com/alibaba-damo-academy/ViSD-Boost/blob/HEAD/lavis/utils/model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"41a29c1c6528fdf6"}},{"code_sha256_prefix":"bde1c545f0335029","entry":"VectorQuantizerEMA","repo":"alibaba-damo-academy/ViSD-Boost","repo_kind":"found_in_text","path":"lavis/utils/model.py","file_url":"https://github.com/alibaba-damo-academy/ViSD-Boost/blob/HEAD/lavis/utils/model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"bde1c545f0335029"}},{"code_sha256_prefix":"df8b11168530146b","entry":"ViT_VAE","repo":"alibaba-damo-academy/ViSD-Boost","repo_kind":"found_in_text","path":"lavis/utils/model.py","file_url":"https://github.com/alibaba-damo-academy/ViSD-Boost/blob/HEAD/lavis/utils/model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"df8b11168530146b"}},{"code_sha256_prefix":"036b339c2e2e5e68","entry":"ViT_VAE_Decoder","repo":"alibaba-damo-academy/ViSD-Boost","repo_kind":"found_in_text","path":"lavis/utils/model.py","file_url":"https://github.com/alibaba-damo-academy/ViSD-Boost/blob/HEAD/lavis/utils/model.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"036b339c2e2e5e68"}},{"code_sha256_prefix":"1bc53a3aa1e759b9","entry":"VISDBOOST","repo":"alibaba-damo-academy/ViSD-Boost","repo_kind":"found_in_text","path":"lavis/utils/model.py","file_url":"https://github.com/alibaba-damo-academy/ViSD-Boost/blob/HEAD/lavis/utils/model.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"1bc53a3aa1e759b9"}}]},"arxiv_metadata":{"licence":"arXiv metadata, CC0 1.0 (https://info.arxiv.org/help/license)","fields":["title","abstract","authors","date"],"primary_category":"eess.IV","source":"arxiv_api"},"syntology_extracted_results":null}