{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/inversion-free-image-editing-with-natural","title":"Inversion-Free Image Editing with Natural Language","arxiv_id":"2312.04965","date":"2023-12-07","proceeding":null,"authors":["Sihan Xu","Yidong Huang","Jiayi Pan","Ziqiao Ma","Joyce Chai"],"abstract":"Despite recent advances in inversion-based editing, text-guided image manipulation remains challenging for diffusion models. The primary bottlenecks include 1) the time-consuming nature of the inversion process; 2) the struggle to balance consistency with accuracy; 3) the lack of compatibility with efficient consistency sampling methods used in consistency models. To address the above issues, we start by asking ourselves if the inversion process can be eliminated for editing. We show that when the initial sample is known, a special variance schedule reduces the denoising step to the same form as the multi-step consistency sampling. We name this Denoising Diffusion Consistent Model (DDCM), and note that it implies a virtual inversion strategy without explicit inversion in sampling. We further unify the attention control mechanisms in a tuning-free framework for text-guided editing. Combining them, we present inversion-free editing (InfEdit), which allows for consistent and faithful editing for both rigid and non-rigid semantic changes, catering to intricate modifications without compromising on the image's integrity and explicit inversion. Through extensive experiments, InfEdit shows strong performance in various editing tasks and also maintains a seamless workflow (less than 3 seconds on one single A40), demonstrating the potential for real-time applications. Project Page: https://sled-group.github.io/InfEdit/","url_abs":"https://arxiv.org/abs/2312.04965v1","url_pdf":"https://arxiv.org/pdf/2312.04965v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"inversion-free-image-editing-with-natural","repo_url":"https://github.com/sled-group/InfEdit","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-manipulation","task_name":"Image Manipulation"},{"task_slug":"text-based-image-editing","task_name":"Text-based Image Editing"}],"methods":[{"method_slug":"diffusion","method_name":"Diffusion"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/text-based-image-editing-on-pie-bench","task":"Text-based Image Editing","dataset":"PIE-Bench","model":"Virtual Inversion+Unified Attention Control+LCM","rank_in_archive_order":2,"of":18,"metrics":{"Background LPIPS":"47.58","Background PSNR":"28.51","CLIPSIM":"25.03","Structure Distance":"13.78"},"uses_additional_data":false},{"leaderboard":"/sota/text-based-image-editing-on-pie-bench","task":"Text-based Image Editing","dataset":"PIE-Bench","model":"Virtual Inversion+Prompt-to-Prompt","rank_in_archive_order":4,"of":18,"metrics":{"Background LPIPS":"47.98","Background PSNR":"27.52","CLIPSIM":"24.89","Structure Distance":"14.22"},"uses_additional_data":false},{"leaderboard":"/sota/text-based-image-editing-on-pie-bench","task":"Text-based Image Editing","dataset":"PIE-Bench","model":"Virtual Inversion+Prompt-to-Prompt+LCM","rank_in_archive_order":7,"of":18,"metrics":{"Background LPIPS":"55.85","Background PSNR":"26.64","CLIPSIM":"24.57","Structure Distance":"15.61"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2312.04965","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2312.04965"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/sled-group/InfEdit","reach":null}],"summary":{"ran_honours":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"9041992d3fb5205d","entry":"replace_nsfw_images","repo":"sled-group/InfEdit","repo_kind":"official","path":"app_infedit.py","file_url":"https://github.com/sled-group/InfEdit/blob/HEAD/app_infedit.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"9041992d3fb5205d"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}