{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/autosam-adapting-sam-to-medical-images-by","title":"AutoSAM: Adapting SAM to Medical Images by Overloading the Prompt Encoder","arxiv_id":"2306.06370","date":"2023-06-10","proceeding":null,"authors":["Tal Shaharabany","Aviad Dahan","Raja Giryes","Lior Wolf"],"abstract":"The recently introduced Segment Anything Model (SAM) combines a clever architecture and large quantities of training data to obtain remarkable image segmentation capabilities. However, it fails to reproduce such results for Out-Of-Distribution (OOD) domains such as medical images. Moreover, while SAM is conditioned on either a mask or a set of points, it may be desirable to have a fully automatic solution. In this work, we replace SAM's conditioning with an encoder that operates on the same input image. By adding this encoder and without further fine-tuning SAM, we obtain state-of-the-art results on multiple medical images and video benchmarks. This new encoder is trained via gradients provided by a frozen SAM. For inspecting the knowledge within it, and providing a lightweight segmentation solution, we also learn to decode it into a mask by a shallow deconvolution network.","url_abs":"https://arxiv.org/abs/2306.06370v1","url_pdf":"https://arxiv.org/pdf/2306.06370v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"image-segmentation","task_name":"Image Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"video-polyp-segmentation","task_name":"Video Polyp Segmentation"}],"methods":[{"method_slug":"sam","method_name":"SAM"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-polyp-segmentation-on-sun-seg-easy","task":"Video Polyp Segmentation","dataset":"SUN-SEG-Easy (Unseen)","model":"AutoSAM","rank_in_archive_order":5,"of":18,"metrics":{"Dice":"0.753","S measure":"0.815","Sensitivity":"0.672","mean E-measure":"0.855","mean F-measure":"0.774","weighted F-measure":"0.716"},"uses_additional_data":false},{"leaderboard":"/sota/video-polyp-segmentation-on-sun-seg-hard","task":"Video Polyp Segmentation","dataset":"SUN-SEG-Hard (Unseen)","model":"AutoSAM","rank_in_archive_order":4,"of":18,"metrics":{"Dice":"0.759","S-Measure":"0.822","Sensitivity":"0.726","mean E-measure":"0.866","mean F-measure":"0.764","weighted F-measure":"0.714"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2306.06370","atlas_url":"https://app.syntology.ai/?focus=2306.06370","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}