{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/superedit-rectifying-and-facilitating","title":"SuperEdit: Rectifying and Facilitating Supervision for Instruction-Based Image Editing","arxiv_id":"2505.02370","date":"2025-05-05","proceeding":null,"authors":["Ming Li","Xin Gu","Fan Chen","Xiaoying Xing","Longyin Wen","Chen Chen","Sijie Zhu"],"abstract":"Due to the challenges of manually collecting accurate editing data, existing datasets are typically constructed using various automated methods, leading to noisy supervision signals caused by the mismatch between editing instructions and original-edited image pairs. Recent efforts attempt to improve editing models through generating higher-quality edited images, pre-training on recognition tasks, or introducing vision-language models (VLMs) but fail to resolve this fundamental issue. In this paper, we offer a novel solution by constructing more effective editing instructions for given image pairs. This includes rectifying the editing instructions to better align with the original-edited image pairs and using contrastive editing instructions to further enhance their effectiveness. Specifically, we find that editing models exhibit specific generation attributes at different inference steps, independent of the text. Based on these prior attributes, we define a unified guide for VLMs to rectify editing instructions. However, there are some challenging editing scenarios that cannot be resolved solely with rectified instructions. To this end, we further construct contrastive supervision signals with positive and negative instructions and introduce them into the model training using triplet loss, thereby further facilitating supervision effectiveness. Our method does not require the VLM modules or pre-training tasks used in previous work, offering a more direct and efficient way to provide better supervision signals, and providing a novel, simple, and effective solution for instruction-based image editing. Results on multiple benchmarks demonstrate that our method significantly outperforms existing approaches. Compared with previous SOTA SmartEdit, we achieve 9.19% improvements on the Real-Edit benchmark with 30x less training data and 13x smaller model size.","url_abs":"https://arxiv.org/abs/2505.02370v1","url_pdf":"https://arxiv.org/pdf/2505.02370v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"superedit-rectifying-and-facilitating","repo_url":"https://github.com/bytedance/superedit","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":null,"task_name":"Triplet"}],"methods":[{"method_slug":"align","method_name":"ALIGN"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2505.02370","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2505.02370"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/bytedance/superedit","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran_draft_wrong":1,"unverified":9},"by_repo_kind":{"official":{"samples":10,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"53d425eb1278ddda","entry":"load_json","repo":"bytedance/superedit","repo_kind":"official","path":"eval/eval_instructpix2pix.py","file_url":"https://github.com/bytedance/superedit/blob/HEAD/eval/eval_instructpix2pix.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"53d425eb1278ddda"}},{"code_sha256_prefix":"db24600360c53d13","entry":"call_azure_gpt4v","repo":"bytedance/superedit","repo_kind":"official","path":"eval/eval_following.py","file_url":"https://github.com/bytedance/superedit/blob/HEAD/eval/eval_following.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"db24600360c53d13"}},{"code_sha256_prefix":"025373d0fac57f74","entry":"call_azure_gpt4v","repo":"bytedance/superedit","repo_kind":"official","path":"eval/eval_keep_detail.py","file_url":"https://github.com/bytedance/superedit/blob/HEAD/eval/eval_keep_detail.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"025373d0fac57f74"}},{"code_sha256_prefix":"ddb9686bf6284670","entry":"call_azure_gpt4v","repo":"bytedance/superedit","repo_kind":"official","path":"eval/eval_quality.py","file_url":"https://github.com/bytedance/superedit/blob/HEAD/eval/eval_quality.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"ddb9686bf6284670"}},{"code_sha256_prefix":"cb354883154c4397","entry":"convert_to_np","repo":"bytedance/superedit","repo_kind":"official","path":"superedit/instruct_pix2pix/train_sd15.py","file_url":"https://github.com/bytedance/superedit/blob/HEAD/superedit/instruct_pix2pix/train_sd15.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"cb354883154c4397"}},{"code_sha256_prefix":"d5288b587cd15fcf","entry":"download_image","repo":"bytedance/superedit","repo_kind":"official","path":"superedit/instruct_pix2pix/train_sd15.py","file_url":"https://github.com/bytedance/superedit/blob/HEAD/superedit/instruct_pix2pix/train_sd15.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"d5288b587cd15fcf"}},{"code_sha256_prefix":"ea7b78776c588d7f","entry":"extract_score","repo":"bytedance/superedit","repo_kind":"official","path":"eval/eval_instructpix2pix.py","file_url":"https://github.com/bytedance/superedit/blob/HEAD/eval/eval_instructpix2pix.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"ea7b78776c588d7f"}},{"code_sha256_prefix":"82a7d117ac0f3760","entry":"load_local_image","repo":"bytedance/superedit","repo_kind":"official","path":"superedit/instruct_pix2pix/train_sd15.py","file_url":"https://github.com/bytedance/superedit/blob/HEAD/superedit/instruct_pix2pix/train_sd15.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"82a7d117ac0f3760"}},{"code_sha256_prefix":"8242d5a10628e9ac","entry":"read_base64_img","repo":"bytedance/superedit","repo_kind":"official","path":"eval/eval_following.py","file_url":"https://github.com/bytedance/superedit/blob/HEAD/eval/eval_following.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"8242d5a10628e9ac"}},{"code_sha256_prefix":"9b82aa92f5718a91","entry":"read_base64_img","repo":"bytedance/superedit","repo_kind":"official","path":"eval/eval_quality.py","file_url":"https://github.com/bytedance/superedit/blob/HEAD/eval/eval_quality.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"9b82aa92f5718a91"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}