{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/towards-open-world-text-guided-face-image","title":"Towards Open-World Text-Guided Face Image Generation and Manipulation","arxiv_id":"2104.08910","date":"2021-04-18","proceeding":null,"authors":["Weihao Xia","Yujiu Yang","Jing-Hao Xue","Baoyuan Wu"],"abstract":"The existing text-guided image synthesis methods can only produce limited quality results with at most \\mbox{$\\text{256}^2$} resolution and the textual instructions are constrained in a small Corpus. In this work, we propose a unified framework for both face image generation and manipulation that produces diverse and high-quality images with an unprecedented resolution at 1024 from multimodal inputs. More importantly, our method supports open-world scenarios, including both image and text, without any re-training, fine-tuning, or post-processing. To be specific, we propose a brand new paradigm of text-guided image generation and manipulation based on the superior characteristics of a pretrained GAN model. Our proposed paradigm includes two novel strategies. The first strategy is to train a text encoder to obtain latent codes that align with the hierarchically semantic of the aforementioned pretrained GAN model. The second strategy is to directly optimize the latent codes in the latent space of the pretrained GAN model with guidance from a pretrained language model. The latent codes can be randomly sampled from a prior distribution or inverted from a given image, which provides inherent supports for both image generation and manipulation from multi-modal inputs, such as sketches or semantic labels, with textual guidance. To facilitate text-guided multi-modal synthesis, we propose the Multi-Modal CelebA-HQ, a large-scale dataset consisting of real face images and corresponding semantic segmentation map, sketch, and textual descriptions. Extensive experiments on the introduced dataset demonstrate the superior performance of our proposed method. Code and data are available at https://github.com/weihaox/TediGAN.","url_abs":"https://arxiv.org/abs/2104.08910v1","url_pdf":"https://arxiv.org/pdf/2104.08910v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"towards-open-world-text-guided-face-image","repo_url":"https://github.com/weihaox/TediGAN","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"towards-open-world-text-guided-face-image","repo_url":"https://github.com/IIGROUP/TediGAN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"image-generation","task_name":"Image Generation"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"text-to-image-generation","task_name":"Text-to-Image Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/text-to-image-generation-on-multi-modal","task":"Text-to-Image Generation","dataset":"Multi-Modal-CelebA-HQ","model":"TediGAN-B","rank_in_archive_order":5,"of":10,"metrics":{"Acc":"20.4","FID":"101.42","LPIPS":"0.461","Real":"21.0"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2104.08910","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2104.08910"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/weihaox/TediGAN","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/IIGROUP/TediGAN","reach":null}],"summary":{"ran_draft_wrong":2,"unverified":3},"by_repo_kind":{"official":{"samples":4,"ran":1,"repositories":1},"listed":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"ff0d399a0317c249","entry":"get_temp_logger","repo":"weihaox/TediGAN","repo_kind":"official","path":"base/models/base_module.py","file_url":"https://github.com/weihaox/TediGAN/blob/HEAD/base/models/base_module.py","link_basis":"plan_row","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"ff0d399a0317c249"}},{"code_sha256_prefix":"4fa4fd0e38079888","entry":"run_on_batch","repo":"IIGROUP/TediGAN","repo_kind":"listed","path":"ext/inference.py","file_url":"https://github.com/IIGROUP/TediGAN/blob/HEAD/ext/inference.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4fa4fd0e38079888"}},{"code_sha256_prefix":"6bc5a9d2ae904282","entry":"build_encoder","repo":"weihaox/TediGAN","repo_kind":"official","path":"base/models/helper.py","file_url":"https://github.com/weihaox/TediGAN/blob/HEAD/base/models/helper.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6bc5a9d2ae904282"}},{"code_sha256_prefix":"01ad09c8b01d79ee","entry":"build_generator","repo":"weihaox/TediGAN","repo_kind":"official","path":"base/models/helper.py","file_url":"https://github.com/weihaox/TediGAN/blob/HEAD/base/models/helper.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"01ad09c8b01d79ee"}},{"code_sha256_prefix":"5eb9fdaa872d49a6","entry":"get_weight_path","repo":"weihaox/TediGAN","repo_kind":"official","path":"base/models/model_settings.py","file_url":"https://github.com/weihaox/TediGAN/blob/HEAD/base/models/model_settings.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"5eb9fdaa872d49a6"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}