{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/delving-stylegan-inversion-for-image-editing","title":"Delving StyleGAN Inversion for Image Editing: A Foundation Latent Space Viewpoint","arxiv_id":"2211.11448","date":"2022-11-21","proceeding":"CVPR 2023 1","authors":["Hongyu Liu","Yibing Song","Qifeng Chen"],"abstract":"GAN inversion and editing via StyleGAN maps an input image into the embedding spaces ($\\mathcal{W}$, $\\mathcal{W^+}$, and $\\mathcal{F}$) to simultaneously maintain image fidelity and meaningful manipulation. From latent space $\\mathcal{W}$ to extended latent space $\\mathcal{W^+}$ to feature space $\\mathcal{F}$ in StyleGAN, the editability of GAN inversion decreases while its reconstruction quality increases. Recent GAN inversion methods typically explore $\\mathcal{W^+}$ and $\\mathcal{F}$ rather than $\\mathcal{W}$ to improve reconstruction fidelity while maintaining editability. As $\\mathcal{W^+}$ and $\\mathcal{F}$ are derived from $\\mathcal{W}$ that is essentially the foundation latent space of StyleGAN, these GAN inversion methods focusing on $\\mathcal{W^+}$ and $\\mathcal{F}$ spaces could be improved by stepping back to $\\mathcal{W}$. In this work, we propose to first obtain the precise latent code in foundation latent space $\\mathcal{W}$. We introduce contrastive learning to align $\\mathcal{W}$ and the image space for precise latent code discovery. %The obtaining process is by using contrastive learning to align $\\mathcal{W}$ and the image space. Then, we leverage a cross-attention encoder to transform the obtained latent code in $\\mathcal{W}$ into $\\mathcal{W^+}$ and $\\mathcal{F}$, accordingly. Our experiments show that our exploration of the foundation latent space $\\mathcal{W}$ improves the representation ability of latent codes in $\\mathcal{W^+}$ and features in $\\mathcal{F}$, which yields state-of-the-art reconstruction fidelity and editability results on the standard benchmarks. Project page: https://kumapowerliu.github.io/CLCAE.","url_abs":"https://arxiv.org/abs/2211.11448v3","url_pdf":"https://arxiv.org/pdf/2211.11448v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"delving-stylegan-inversion-for-image-editing","repo_url":"https://github.com/kumapowerliu/clcae","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"contrastive-learning","task_name":"Contrastive Learning"}],"methods":[{"method_slug":"align","method_name":"ALIGN"},{"method_slug":"adaptive-instance-normalization","method_name":"Adaptive Instance Normalization"},{"method_slug":"contrastive-learning","method_name":"Contrastive Learning"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"feedforward-network","method_name":"Feedforward Network"},{"method_slug":"r1-regularization","method_name":"R1 Regularization"},{"method_slug":"stylegan","method_name":"StyleGAN"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2211.11448","atlas_url":"https://app.syntology.ai/?focus=2211.11448","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}