{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/rs-agent-automating-remote-sensing-tasks","title":"RS-Agent: Automating Remote Sensing Tasks through Intelligent Agent","arxiv_id":"2406.07089","date":"2024-06-11","proceeding":null,"authors":["Wenjia Xu","Zijian Yu","Boyang Mu","Zhiwei Wei","Yuanben Zhang","Guangzuo Li","Mugen Peng"],"abstract":"The unprecedented advancements in Multimodal Large Language Models (MLLMs) have demonstrated strong potential in interacting with humans through both language and visual inputs to perform downstream tasks such as visual question answering and scene understanding. However, these models are constrained to basic instruction-following or descriptive tasks, facing challenges in complex real-world remote sensing applications that require specialized tools and knowledge. To address these limitations, we propose RS-Agent, an AI agent designed to interact with human users and autonomously leverage specialized models to address the demands of real-world remote sensing applications. RS-Agent integrates four key components: a Central Controller based on large language models, a dynamic toolkit for tool execution, a Solution Space for task-specific expert guidance, and a Knowledge Space for domain-level reasoning, enabling it to interpret user queries and orchestrate tools for accurate remote sensing task. We introduce two novel mechanisms: Task-Aware Retrieval, which improves tool selection accuracy through expert-guided planning, and DualRAG, a retrieval-augmented generation method that enhances knowledge relevance through weighted, dual-path retrieval. RS-Agent supports flexible integration of new tools and is compatible with both open-source and proprietary LLMs. Extensive experiments across 9 datasets and 18 remote sensing tasks demonstrate that RS-Agent significantly outperforms state-of-the-art MLLMs, achieving over 95% task planning accuracy and delivering superior performance in tasks such as scene classification, object counting, and remote sensing visual question answering. Our work presents RS-Agent as a robust and extensible framework for advancing intelligent automation in remote sensing analysis.","url_abs":"https://arxiv.org/abs/2406.07089v2","url_pdf":"https://arxiv.org/pdf/2406.07089v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"rs-agent-automating-remote-sensing-tasks","repo_url":"https://github.com/intellisensing/rs-agent","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"ai-agent","task_name":"AI Agent"},{"task_slug":"descriptive","task_name":"Descriptive"},{"task_slug":"instruction-following","task_name":"Instruction Following"},{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"large-language-model","task_name":"Large Language Model"},{"task_slug":"object-counting","task_name":"Object Counting"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"retrieval-augmented-generation","task_name":"Retrieval-augmented Generation"},{"task_slug":"scene-classification","task_name":"Scene Classification"},{"task_slug":"scene-understanding","task_name":"Scene Understanding"},{"task_slug":"task-planning","task_name":"Task Planning"},{"task_slug":"visual-question-answering-1","task_name":"Visual Question Answering"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2406.07089","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2406.07089"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/intellisensing/rs-agent","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran":4,"unverified":1},"by_repo_kind":{"official":{"samples":5,"ran":4,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"746af84399f80aba","entry":"extract_first_tool","repo":"intellisensing/rs-agent","repo_kind":"official","path":"benchmarks/planning/score.py","file_url":"https://github.com/intellisensing/rs-agent/blob/HEAD/benchmarks/planning/score.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"746af84399f80aba"}},{"code_sha256_prefix":"1ec7ac3eddfe5bff","entry":"load_config","repo":"intellisensing/rs-agent","repo_kind":"official","path":"rs_agent/config.py","file_url":"https://github.com/intellisensing/rs-agent/blob/HEAD/rs_agent/config.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"1ec7ac3eddfe5bff"}},{"code_sha256_prefix":"0456685d7dc9d641","entry":"normalize_task_type","repo":"intellisensing/rs-agent","repo_kind":"official","path":"benchmarks/planning/score.py","file_url":"https://github.com/intellisensing/rs-agent/blob/HEAD/benchmarks/planning/score.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"0456685d7dc9d641"}},{"code_sha256_prefix":"c022f7fe5e2beab3","entry":"resolve_path","repo":"intellisensing/rs-agent","repo_kind":"official","path":"rs_agent/config.py","file_url":"https://github.com/intellisensing/rs-agent/blob/HEAD/rs_agent/config.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"c022f7fe5e2beab3"}},{"code_sha256_prefix":"082fb94fcb4b1d9f","entry":"extract_all_tools","repo":"intellisensing/rs-agent","repo_kind":"official","path":"benchmarks/planning/score.py","file_url":"https://github.com/intellisensing/rs-agent/blob/HEAD/benchmarks/planning/score.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"082fb94fcb4b1d9f"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}