{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lrsclip-a-vision-language-foundation-model","title":"LRSCLIP: A Vision-Language Foundation Model for Aligning Remote Sensing Image with Longer Text","arxiv_id":"2503.19311","date":"2025-03-25","proceeding":null,"authors":["Weizhi Chen","Jingbo Chen","Yupeng Deng","Jiansheng Chen","Yuman Feng","Zhihao Xi","Diyou Liu","Kai Li","Yu Meng"],"abstract":"This study addresses the technical bottlenecks in handling long text and the \"hallucination\" issue caused by insufficient short text information in remote sensing vision-language foundation models (VLFM). We propose a novel vision-language foundation model, LRSCLIP, and a multimodal dataset, LRS2M. The main contributions are as follows: (1) By integrating multi-source remote sensing data and adopting a large language model labeling strategy, we construct the LRS2M dataset, which contains 2 million image-text pairs, providing both short and long texts for the first time, thus solving the problem of semantic granularity limitations in existing datasets; (2) The design of the LRSCLIP architecture based on Long-CLIP's KPS module, which extends CLIP's text processing capacity and achieves fine-grained cross-modal feature alignment through a dual-text loss weighting mechanism. Experimental results show that LRSCLIP improves retrieval accuracy by 10\\%-20\\% over the Long-CLIP baseline in the zero-shot long-text cross-modal retrieval task. For the zero-shot short-text cross-modal retrieval task, LRSCLIP achieves improvements over the current best model, GeoRSCLIP, with increases of 0.17\\%, 0.67\\%, and 0.92\\% in Text to Image R@1, Image to Text R@1, and mR on RSITMD, respectively, and 0.04\\%, 2.93\\%, and 1.28\\% on RSICD. In the zero-shot image classification task (average accuracy=75.75\\%) and semantic localization task (Rmi=0.7653), LRSCLIP achieves state-of-the-art performance. These results validate the dual advantages of fine-grained semantic understanding and global feature matching in LRSCLIP. This work provides a new benchmark model and data support for remote sensing multimodal learning. The related code has been open source and is available at https://github.com/MitsuiChen14/LRSCLIP.","url_abs":"https://arxiv.org/abs/2503.19311v1","url_pdf":"https://arxiv.org/pdf/2503.19311v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"lrsclip-a-vision-language-foundation-model","repo_url":"https://github.com/mitsuichen14/lrsclip","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"cross-modal-retrieval","task_name":"Cross-Modal Retrieval"},{"task_slug":"hallucination","task_name":"Hallucination"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"image-to-text","task_name":"Image to text"},{"task_slug":"large-language-model","task_name":"Large Language Model"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"zero-shot-image-classification","task_name":"Zero-Shot Image Classification"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2503.19311","atlas_url":"https://app.syntology.ai/?focus=2503.19311","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}