{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/self-prompting-analogical-reasoning-for-uav","title":"self-prompting analogical reasoning for uav object detection","arxiv_id":null,"date":"2025-04-11","proceeding":"Proceedings of the AAAI Conference on Artificial Intelligence 2025 4","authors":["Nianxin Li","Mao Ye","Lihua Zhou","Song Tang","Yan Gan","Zizhuo Liang","Xiatian Zhu"],"abstract":"Unmanned Aerial Vehicle Object Detection \r\n(UAVOD)\r\npresents unique challenges due to varying altitudes, dynamic\r\nbackgrounds, and the small size of objects. Traditional de-\r\ntection methods often struggle with these challenges, as\r\nthey typically rely on visual features only and fail to ex-\r\ntract the semantic relations between the objects. To address\r\nthese limitations,we propose a novel approach named Self-\r\nPrompting Analogical Reasoning (SPAR). Ourmethod uti-\r\nlizes the vision-languagemodel (CLIP) to generate context-\r\naware prompts based on image features, providing rich se-\r\nmantic information that guides analogical reasoning. SPAR\r\nincludes two main modules: self-prompting and analogi-\r\ncal reasoning. Self-prompting module based on learnable\r\ndescription and CLIP-text encoder generates context-aware\r\nprompt by combining specific image feature; then an object-\r\nness prompt scoremap is produced by computing the simi-\r\nlaritybetweenpixel-level features andcontext-awareprompt.\r\nWiththis scoremap,multi-scale image features are enhanced\r\nand pixel-level features are chosen for graph construction.\r\nWhile for analogical reasoningmodule, graph nodes consist\r\nof category-level prompt nodes and pixel-level image feature\r\nnodes.Analogical inference is based on graph convolution.\r\nUnder the guidance of category-level nodes, different-scale\r\nobject features have been enhanced, which helps achieve\r\nmore accuratedetectionof challengingobjects.Extensive ex-\r\nperiments illustrate that SPARoutperforms traditionalmeth-\r\nods,offeringamorerobustandaccuratesolutionforUAVOD.","url_abs":"https://ojs.aaai.org/index.php/AAAI/article/view/34026","url_pdf":"https://ojs.aaai.org/index.php/AAAI/article/view/34026/36181","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"self-prompting-analogical-reasoning-for-uav","repo_url":"https://github.com/lnxwow/Analogical-Reasoning","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"object-detection-in-aerial-images","task_name":"Object Detection In Aerial Images"},{"task_slug":"small-object-detection","task_name":"Small Object Detection"},{"task_slug":"graph-construction","task_name":"graph construction"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[{"method_slug":"clip","method_name":"CLIP"},{"method_slug":"yolov8","method_name":"YOLOv8"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}