{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mmm-rs-a-multi-modal-multi-gsd-multi-scene","title":"MMM-RS: A Multi-modal, Multi-GSD, Multi-scene Remote Sensing Dataset and Benchmark for Text-to-Image Generation","arxiv_id":"2410.22362","date":"2024-10-26","proceeding":null,"authors":["Jialin Luo","Yuanzhi Wang","Ziqi Gu","Yide Qiu","Shuaizhen Yao","Fuyun Wang","Chunyan Xu","Wenhua Zhang","Dan Wang","Zhen Cui"],"abstract":"Recently, the diffusion-based generative paradigm has achieved impressive general image generation capabilities with text prompts due to its accurate distribution modeling and stable training process. However, generating diverse remote sensing (RS) images that are tremendously different from general images in terms of scale and perspective remains a formidable challenge due to the lack of a comprehensive remote sensing image generation dataset with various modalities, ground sample distances (GSD), and scenes. In this paper, we propose a Multi-modal, Multi-GSD, Multi-scene Remote Sensing (MMM-RS) dataset and benchmark for text-to-image generation in diverse remote sensing scenarios. Specifically, we first collect nine publicly available RS datasets and conduct standardization for all samples. To bridge RS images to textual semantic information, we utilize a large-scale pretrained vision-language model to automatically output text prompts and perform hand-crafted rectification, resulting in information-rich text-image pairs (including multi-modal images). In particular, we design some methods to obtain the images with different GSD and various environments (e.g., low-light, foggy) in a single sample. With extensive manual screening and refining annotations, we ultimately obtain a MMM-RS dataset that comprises approximately 2.1 million text-image pairs. Extensive experimental results verify that our proposed MMM-RS dataset allows off-the-shelf diffusion models to generate diverse RS images across various modalities, scenes, weather conditions, and GSD. The dataset is available at https://github.com/ljl5261/MMM-RS.","url_abs":"https://arxiv.org/abs/2410.22362v1","url_pdf":"https://arxiv.org/pdf/2410.22362v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mmm-rs-a-multi-modal-multi-gsd-multi-scene","repo_url":"https://github.com/ljl5261/mmm-rs","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"text-to-image-generation-1","task_name":"Text to Image Generation"},{"task_slug":"text-to-image-generation","task_name":"Text-to-Image Generation"}],"methods":[{"method_slug":"diffusion","method_name":"Diffusion"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2410.22362","atlas_url":"https://app.syntology.ai/?focus=2410.22362","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2410.22362"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"deterministic:regex_extraction","url":"https://github.com/ljl5261/MMM-RS","reach":{"status":"ok"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/ljl5261/mmm-rs","reach":{"status":"ok"}}],"summary":{"unverified":5},"by_repo_kind":{"official":{"samples":5,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":5,"samples":[{"code_sha256_prefix":"5e8db6d8b9a566fb","entry":"center_crop","repo":"ljl5261/MMM-RS","repo_kind":"official","path":"Standardized_processing/HRSC2016.py","file_url":"https://github.com/ljl5261/MMM-RS/blob/HEAD/Standardized_processing/HRSC2016.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5e8db6d8b9a566fb"}},{"code_sha256_prefix":"853f808100e34cb3","entry":"center_crop","repo":"ljl5261/MMM-RS","repo_kind":"official","path":"Standardized_processing/TGRS-HRRSD.py","file_url":"https://github.com/ljl5261/MMM-RS/blob/HEAD/Standardized_processing/TGRS-HRRSD.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"853f808100e34cb3"}},{"code_sha256_prefix":"feeab8376692a58d","entry":"center_crop","repo":"ljl5261/MMM-RS","repo_kind":"official","path":"Standardized_processing/fMoW.py","file_url":"https://github.com/ljl5261/MMM-RS/blob/HEAD/Standardized_processing/fMoW.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"feeab8376692a58d"}},{"code_sha256_prefix":"4d359f2b3b1ef109","entry":"crop_images","repo":"ljl5261/MMM-RS","repo_kind":"official","path":"Standardized_processing/fMoW.py","file_url":"https://github.com/ljl5261/MMM-RS/blob/HEAD/Standardized_processing/fMoW.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"4d359f2b3b1ef109"}},{"code_sha256_prefix":"3a0e0bf0aa97110e","entry":"super_resolve_image_batch","repo":"ljl5261/MMM-RS","repo_kind":"official","path":"Standardized_processing/SEN1-2.py","file_url":"https://github.com/ljl5261/MMM-RS/blob/HEAD/Standardized_processing/SEN1-2.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"3a0e0bf0aa97110e"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}