{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/chatearthnet-a-global-scale-high-quality","title":"ChatEarthNet: A Global-Scale Image-Text Dataset Empowering Vision-Language Geo-Foundation Models","arxiv_id":"2402.11325","date":"2024-02-17","proceeding":null,"authors":["Zhenghang Yuan","Zhitong Xiong","Lichao Mou","Xiao Xiang Zhu"],"abstract":"An in-depth comprehension of global land cover is essential in Earth observation, forming the foundation for a multitude of applications. Although remote sensing technology has advanced rapidly, leading to a proliferation of satellite imagery, the inherent complexity of these images often makes them difficult for non-expert users to understand. Natural language, as a carrier of human knowledge, can be a bridge between common users and complicated satellite imagery. In this context, we introduce a global-scale, high-quality image-text dataset for remote sensing, providing natural language descriptions for Sentinel-2 data to facilitate the understanding of satellite imagery for common users. Specifically, we utilize Sentinel-2 data for its global coverage as the foundational image source, employing semantic segmentation labels from the European Space Agency's (ESA) WorldCover project to enrich the descriptions of land covers. By conducting in-depth semantic analysis, we formulate detailed prompts to elicit rich descriptions from ChatGPT. To enhance the dataset's quality, we introduce the manual verification process. This step involves manual inspection and correction to refine the dataset, thus significantly improving its accuracy and quality. Finally, we offer the community ChatEarthNet, a large-scale image-text dataset characterized by global coverage, high quality, wide-ranging diversity, and detailed descriptions. ChatEarthNet consists of 163,488 image-text pairs with captions generated by ChatGPT-3.5 and an additional 10,000 image-text pairs with captions generated by ChatGPT-4V(ision). This dataset has significant potential for training vision-language geo-foundation models and evaluating large vision-language models for remote sensing. The dataset will be made publicly available.","url_abs":"https://arxiv.org/abs/2402.11325v2","url_pdf":"https://arxiv.org/pdf/2402.11325v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"chatearthnet-a-global-scale-high-quality","repo_url":"https://github.com/zhu-xlab/ChatEarthNet","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"earth-observation","task_name":"Earth Observation"},{"task_slug":"image-captioning","task_name":"Image Captioning"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[{"slug":"chatearthnet","name":"ChatEarthNet","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2402.11325","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2402.11325"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/zhu-xlab/ChatEarthNet","reach":{"status":"ok"}}],"summary":{"ran":6},"by_repo_kind":{"official":{"samples":6,"ran":6,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":6,"samples":[{"code_sha256_prefix":"feb0491542bc3631","entry":"analyze_segmentation_map","repo":"zhu-xlab/ChatEarthNet","repo_kind":"official","path":"generate_description_chatgpt4v.py","file_url":"https://github.com/zhu-xlab/ChatEarthNet/blob/HEAD/generate_description_chatgpt4v.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"feb0491542bc3631"}},{"code_sha256_prefix":"7bb8961562c5043f","entry":"convert_color_map_to_segmentation","repo":"zhu-xlab/ChatEarthNet","repo_kind":"official","path":"generate_description_chatgpt3_5_text.py","file_url":"https://github.com/zhu-xlab/ChatEarthNet/blob/HEAD/generate_description_chatgpt3_5_text.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7bb8961562c5043f"}},{"code_sha256_prefix":"f7cb77470f1615ae","entry":"count_pixel_proportions","repo":"zhu-xlab/ChatEarthNet","repo_kind":"official","path":"generate_description_chatgpt4v.py","file_url":"https://github.com/zhu-xlab/ChatEarthNet/blob/HEAD/generate_description_chatgpt4v.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"f7cb77470f1615ae"}},{"code_sha256_prefix":"a94b22fcb196a7ab","entry":"count_pixel_proportions1","repo":"zhu-xlab/ChatEarthNet","repo_kind":"official","path":"generate_description_chatgpt3_5_text.py","file_url":"https://github.com/zhu-xlab/ChatEarthNet/blob/HEAD/generate_description_chatgpt3_5_text.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a94b22fcb196a7ab"}},{"code_sha256_prefix":"a864fd8808756331","entry":"divide_into_patches","repo":"zhu-xlab/ChatEarthNet","repo_kind":"official","path":"generate_description_chatgpt3_5_text.py","file_url":"https://github.com/zhu-xlab/ChatEarthNet/blob/HEAD/generate_description_chatgpt3_5_text.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"a864fd8808756331"}},{"code_sha256_prefix":"5587fd57c8d667c8","entry":"split_into_patches","repo":"zhu-xlab/ChatEarthNet","repo_kind":"official","path":"generate_description_chatgpt4v.py","file_url":"https://github.com/zhu-xlab/ChatEarthNet/blob/HEAD/generate_description_chatgpt4v.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"5587fd57c8d667c8"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}