{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/rs5m-a-large-scale-vision-language-dataset","title":"RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing","arxiv_id":"2306.11300","date":"2023-06-20","proceeding":null,"authors":["Zilun Zhang","Tiancheng Zhao","Yulong Guo","Jianwei Yin"],"abstract":"Pre-trained Vision-Language Models (VLMs) utilizing extensive image-text paired data have demonstrated unprecedented image-text association capabilities, achieving remarkable results across various downstream tasks. A critical challenge is how to make use of existing large-scale pre-trained VLMs, which are trained on common objects, to perform the domain-specific transfer for accomplishing domain-related downstream tasks. A critical challenge is how to make use of existing large-scale pre-trained VLMs, which are trained on common objects, to perform the domain-specific transfer for accomplishing domain-related downstream tasks. In this paper, we propose a new framework that includes the Domain pre-trained Vision-Language Model (DVLM), bridging the gap between the General Vision-Language Model (GVLM) and domain-specific downstream tasks. Moreover, we present an image-text paired dataset in the field of remote sensing (RS), RS5M, which has 5 million RS images with English descriptions. The dataset is obtained from filtering publicly available image-text paired datasets and captioning label-only RS datasets with pre-trained VLM. These constitute the first large-scale RS image-text paired dataset. Additionally, we fine-tuned the CLIP model and tried several Parameter-Efficient Fine-Tuning methods on RS5M to implement the DVLM. Experimental results show that our proposed dataset is highly effective for various tasks, and our model GeoRSCLIP improves upon the baseline or previous state-of-the-art model by $3\\%\\sim20\\%$ in Zero-shot Classification (ZSC), $3\\%\\sim6\\%$ in Remote Sensing Cross-Modal Text-Image Retrieval (RSCTIR) and $4\\%\\sim5\\%$ in Semantic Localization (SeLo) tasks. Dataset and models have been released in: \\url{https://github.com/om-ai-lab/RS5M}.","url_abs":"https://arxiv.org/abs/2306.11300v5","url_pdf":"https://arxiv.org/pdf/2306.11300v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"rs5m-a-large-scale-vision-language-dataset","repo_url":"https://github.com/om-ai-lab/rs5m","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"cross-modal-retrieval","task_name":"Cross-Modal Retrieval"},{"task_slug":"image-retrieval","task_name":"Image Retrieval"},{"task_slug":"image-to-text-retrieval","task_name":"Image-to-Text Retrieval"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"text-retrieval","task_name":"Text Retrieval"},{"task_slug":"zero-shot-learning","task_name":"Zero-Shot Learning"},{"task_slug":"parameter-efficient-fine-tuning","task_name":"parameter-efficient fine-tuning"},{"task_slug":null,"task_name":"zero-shot-classification"}],"methods":[{"method_slug":"clip","method_name":"CLIP"},{"method_slug":"diffusion","method_name":"Diffusion"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/cross-modal-retrieval-on-rsicd","task":"Cross-Modal Retrieval","dataset":"RSICD","model":"GeoRSCLIP-FT","rank_in_archive_order":2,"of":10,"metrics":{"Image-to-text R@1":"21.13%","Mean Recall":"38.87%","text-to-image R@1":"15.59%"},"uses_additional_data":true},{"leaderboard":"/sota/cross-modal-retrieval-on-rsitmd","task":"Cross-Modal Retrieval","dataset":"RSITMD","model":"GeoRSCLIP-FT","rank_in_archive_order":2,"of":10,"metrics":{"Image-to-text R@1":"32.30%","Mean Recall":"51.81%","text-to-imageR@1":"25.04%"},"uses_additional_data":true},{"leaderboard":"/sota/image-to-text-retrieval-on-rsicd","task":"Image-to-Text Retrieval","dataset":"RSICD","model":"GeoRSCLIP-FT","rank_in_archive_order":1,"of":1,"metrics":{"Image to Text Recall@1":"22.14%"},"uses_additional_data":true},{"leaderboard":"/sota/text-retrieval-on-rsicd","task":"Text Retrieval","dataset":"RSICD","model":"GeoRSCLIP-FT","rank_in_archive_order":1,"of":1,"metrics":{"Recall@1":"15.59%"},"uses_additional_data":true}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2306.11300","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2306.11300"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/om-ai-lab/rs5m","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"unverified":7},"by_repo_kind":{"official":{"samples":7,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"e422ef5eabd3e2d3","entry":"build_ordinary_dataset_dataloader","repo":"om-ai-lab/rs5m","repo_kind":"official","path":"dataloader.py","file_url":"https://github.com/om-ai-lab/rs5m/blob/HEAD/dataloader.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e422ef5eabd3e2d3"}},{"code_sha256_prefix":"0106dca5c51f8f1e","entry":"build_redcaps_url_utc_dict","repo":"om-ai-lab/rs5m","repo_kind":"official","path":"geometa_extraction_tool/redcaps_geometa_caption.py","file_url":"https://github.com/om-ai-lab/rs5m/blob/HEAD/geometa_extraction_tool/redcaps_geometa_caption.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0106dca5c51f8f1e"}},{"code_sha256_prefix":"efe314157da16dc0","entry":"convert_utc_timestamp","repo":"om-ai-lab/rs5m","repo_kind":"official","path":"geometa_extraction_tool/helper.py","file_url":"https://github.com/om-ai-lab/rs5m/blob/HEAD/geometa_extraction_tool/helper.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"efe314157da16dc0"}},{"code_sha256_prefix":"070fc16100b59dfe","entry":"convert_yfcc_timestamp","repo":"om-ai-lab/rs5m","repo_kind":"official","path":"geometa_extraction_tool/helper.py","file_url":"https://github.com/om-ai-lab/rs5m/blob/HEAD/geometa_extraction_tool/helper.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"070fc16100b59dfe"}},{"code_sha256_prefix":"adca0be395aff485","entry":"get_meteor_score","repo":"om-ai-lab/rs5m","repo_kind":"official","path":"blip2_finetune/blip2_peft_inference.py","file_url":"https://github.com/om-ai-lab/rs5m/blob/HEAD/blip2_finetune/blip2_peft_inference.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"adca0be395aff485"}},{"code_sha256_prefix":"1e7ca5b9b019d5bf","entry":"select_tags","repo":"om-ai-lab/rs5m","repo_kind":"official","path":"geometa_extraction_tool/cc3m_geometa_caption.py","file_url":"https://github.com/om-ai-lab/rs5m/blob/HEAD/geometa_extraction_tool/cc3m_geometa_caption.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"1e7ca5b9b019d5bf"}},{"code_sha256_prefix":"52f9aa6edd7ad06f","entry":"split_df","repo":"om-ai-lab/rs5m","repo_kind":"official","path":"geometa_extraction_tool/v35_making_tool_pub11.py","file_url":"https://github.com/om-ai-lab/rs5m/blob/HEAD/geometa_extraction_tool/v35_making_tool_pub11.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"52f9aa6edd7ad06f"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}