{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/dara-domain-and-relation-aware-adapters-make","title":"DARA: Domain- and Relation-aware Adapters Make Parameter-efficient Tuning for Visual Grounding","arxiv_id":"2405.06217","date":"2024-05-10","proceeding":null,"authors":["Ting Liu","Xuyang Liu","Siteng Huang","Honggang Chen","Quanjun Yin","Long Qin","Donglin Wang","Yue Hu"],"abstract":"Visual grounding (VG) is a challenging task to localize an object in an image based on a textual description. Recent surge in the scale of VG models has substantially improved performance, but also introduced a significant burden on computational costs during fine-tuning. In this paper, we explore applying parameter-efficient transfer learning (PETL) to efficiently transfer the pre-trained vision-language knowledge to VG. Specifically, we propose \\textbf{DARA}, a novel PETL method comprising \\underline{\\textbf{D}}omain-aware \\underline{\\textbf{A}}dapters (DA Adapters) and \\underline{\\textbf{R}}elation-aware \\underline{\\textbf{A}}dapters (RA Adapters) for VG. DA Adapters first transfer intra-modality representations to be more fine-grained for the VG domain. Then RA Adapters share weights to bridge the relation between two modalities, improving spatial reasoning. Empirical results on widely-used benchmarks demonstrate that DARA achieves the best accuracy while saving numerous updated parameters compared to the full fine-tuning and other PETL methods. Notably, with only \\textbf{2.13\\%} tunable backbone parameters, DARA improves average accuracy by \\textbf{0.81\\%} across the three benchmarks compared to the baseline model. Our code is available at \\url{https://github.com/liuting20/DARA}.","url_abs":"https://arxiv.org/abs/2405.06217v2","url_pdf":"https://arxiv.org/pdf/2405.06217v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"dara-domain-and-relation-aware-adapters-make","repo_url":"https://github.com/liuting20/dara","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":null,"task_name":"Relation"},{"task_slug":"spatial-reasoning","task_name":"Spatial Reasoning"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"},{"task_slug":"visual-grounding","task_name":"Visual Grounding"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2405.06217","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}