{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/visual-prompting-via-image-inpainting","title":"Visual Prompting via Image Inpainting","arxiv_id":"2209.00647","date":"2022-09-01","proceeding":null,"authors":["Amir Bar","Yossi Gandelsman","Trevor Darrell","Amir Globerson","Alexei A. Efros"],"abstract":"How does one adapt a pre-trained visual model to novel downstream tasks without task-specific finetuning or any model modification? Inspired by prompting in NLP, this paper investigates visual prompting: given input-output image example(s) of a new task at test time and a new input image, the goal is to automatically produce the output image, consistent with the given examples. We show that posing this problem as simple image inpainting - literally just filling in a hole in a concatenated visual prompt image - turns out to be surprisingly effective, provided that the inpainting algorithm has been trained on the right data. We train masked auto-encoders on a new dataset that we curated - 88k unlabeled figures from academic papers sources on Arxiv. We apply visual prompting to these pretrained models and demonstrate results on various downstream image-to-image tasks, including foreground segmentation, single object detection, colorization, edge detection, etc.","url_abs":"https://arxiv.org/abs/2209.00647v1","url_pdf":"https://arxiv.org/pdf/2209.00647v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"visual-prompting-via-image-inpainting","repo_url":"https://github.com/amirbar/visual_prompting","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"colorization","task_name":"Colorization"},{"task_slug":"edge-detection","task_name":"Edge Detection"},{"task_slug":"foreground-segmentation","task_name":"Foreground Segmentation"},{"task_slug":"image-inpainting","task_name":"Image Inpainting"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"personalized-segmentation","task_name":"Personalized Segmentation"},{"task_slug":"visual-prompting","task_name":"Visual Prompting"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[{"method_slug":"pixel-prediction","method_name":"Inpainting"},{"method_slug":"test","method_name":"Test"}],"datasets_introduced":[{"slug":"computer-vision-arxiv-figures","name":"Computer Vision Arxiv Figures","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/personalized-segmentation-on-perseg","task":"Personalized Segmentation","dataset":"PerSeg","model":"Visual Prompting","rank_in_archive_order":5,"of":6,"metrics":{"mIoU":"65.88"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2209.00647","atlas_url":"https://app.syntology.ai/?focus=2209.00647","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}