{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/no-time-to-train-training-free-reference","title":"No time to train! Training-Free Reference-Based Instance Segmentation","arxiv_id":"2507.02798","date":"2025-07-03","proceeding":null,"authors":["Miguel Espinosa","Chenhongyi Yang","Linus Ericsson","Steven McDonagh","Elliot J. Crowley"],"abstract":"The performance of image segmentation models has historically been constrained by the high cost of collecting large-scale annotated data. The Segment Anything Model (SAM) alleviates this original problem through a promptable, semantics-agnostic, segmentation paradigm and yet still requires manual visual-prompts or complex domain-dependent prompt-generation rules to process a new image. Towards reducing this new burden, our work investigates the task of object segmentation when provided with, alternatively, only a small set of reference images. Our key insight is to leverage strong semantic priors, as learned by foundation models, to identify corresponding regions between a reference and a target image. We find that correspondences enable automatic generation of instance-level segmentation masks for downstream tasks and instantiate our ideas via a multi-stage, training-free method incorporating (1) memory bank construction; (2) representation aggregation and (3) semantic-aware feature matching. Our experiments show significant improvements on segmentation metrics, leading to state-of-the-art performance on COCO FSOD (36.8% nAP), PASCAL VOC Few-Shot (71.2% nAP50) and outperforming existing training-free approaches on the Cross-Domain FSOD benchmark (22.4% nAP).","url_abs":"https://arxiv.org/abs/2507.02798v1","url_pdf":"https://arxiv.org/pdf/2507.02798v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"no-time-to-train-training-free-reference","repo_url":"https://github.com/miquel-espinosa/no-time-to-train","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"no-time-to-train-training-free-reference","repo_url":"https://github.com/miquel-espinosa/samantics","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"cross-domain-few-shot-object-detection","task_name":"Cross-Domain Few-Shot Object Detection"},{"task_slug":"few-shot-object-detection","task_name":"Few-Shot Object Detection"},{"task_slug":"image-segmentation","task_name":"Image Segmentation"},{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[{"method_slug":"set","method_name":"SET"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/cross-domain-few-shot-object-detection-on","task":"Cross-Domain Few-Shot Object Detection","dataset":"Artaxor","model":"Training-free(w/o FT)","rank_in_archive_order":8,"of":16,"metrics":{" mAP":"35.0"},"uses_additional_data":false},{"leaderboard":"/sota/cross-domain-few-shot-object-detection-on-1","task":"Cross-Domain Few-Shot Object Detection","dataset":"Clipark1k","model":"Training-free(w/o FT)","rank_in_archive_order":6,"of":10,"metrics":{" mAP":"25.9"},"uses_additional_data":false},{"leaderboard":"/sota/cross-domain-few-shot-object-detection-on-2","task":"Cross-Domain Few-Shot Object Detection","dataset":"DIOR","model":"Training-free(w/o FT)","rank_in_archive_order":11,"of":15,"metrics":{"mAP":"16.4"},"uses_additional_data":false},{"leaderboard":"/sota/cross-domain-few-shot-object-detection-on-3","task":"Cross-Domain Few-Shot Object Detection","dataset":"DeepFish","model":"Training-free(w/o FT)","rank_in_archive_order":3,"of":10,"metrics":{"mAP":"29.6"},"uses_additional_data":false},{"leaderboard":"/sota/cross-domain-few-shot-object-detection-on-neu","task":"Cross-Domain Few-Shot Object Detection","dataset":"NEU-DET","model":"Training-free(w/o FT)","rank_in_archive_order":7,"of":10,"metrics":{"mAP":"5.5"},"uses_additional_data":false},{"leaderboard":"/sota/cross-domain-few-shot-object-detection-on-4","task":"Cross-Domain Few-Shot Object Detection","dataset":"UODD","model":"Training-free(w/o FT)","rank_in_archive_order":7,"of":16,"metrics":{"mAP":"16.0"},"uses_additional_data":false},{"leaderboard":"/sota/few-shot-object-detection-on-ms-coco-1-shot","task":"Few-Shot Object Detection","dataset":"MS-COCO (1-shot)","model":"Training-free","rank_in_archive_order":1,"of":7,"metrics":{"AP":"26.5"},"uses_additional_data":false},{"leaderboard":"/sota/few-shot-object-detection-on-ms-coco-10-shot","task":"Few-Shot Object Detection","dataset":"MS-COCO (10-shot)","model":"Training-free","rank_in_archive_order":1,"of":33,"metrics":{"AP":"36.6"},"uses_additional_data":false},{"leaderboard":"/sota/few-shot-object-detection-on-ms-coco-30-shot","task":"Few-Shot Object Detection","dataset":"MS-COCO (30-shot)","model":"Training-free","rank_in_archive_order":1,"of":25,"metrics":{"AP":"36.8"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}