Browse State-of-the-Art › Zero-Shot Text-to-Image Generation

Zero-Shot Text-to-Image Generation

11 papers with code · 0 benchmarks · 0 datasets archive 2025-07-28

Computer VisionNatural Language Processing

Description from the archive archive 2025-07-28; Papers-with-Code links inside it are rewritten to this site.

Benchmarks archive 2025-07-28

No benchmark for this task in the archive.

Libraries

Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.

Datasets archive 2025-07-28

No dataset record in the archive lists this task.

Subtasks archive 2025-07-28

No subtask under this task in the archive's task tree.

Parent tasks archive 2025-07-28

Most implemented papers archive 2025-07-28

11 shown of 11 papers with code (16 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.

  • 24 Feb 2021 12 repositories listed Syntology ran 7 of 7 samples · 0 unverified · 3 pointer-only (licence)
    Text-to-image generation has traditionally focused on finding better modeling assumptions for training on a fixed dataset.
  • 13 Apr 2022 8 repositories listed Syntology ran 29 of 38 samples · 9 unverified · 1 pointer-only (licence)
    Contrastive models like CLIP have been shown to learn robust representations of images that capture both semantics and style.
  • 26 May 2021 4 repositories listed Syntology ran 4 of 6 samples · 2 unverified
    Text-to-Image generation in the general domain has long been an open problem, which requires both a powerful generative model and cross-modal understanding.
  • 27 Nov 2021 3 repositories listed Syntology ran 4 of 18 samples · 14 unverified
    One of the major challenges in training text-to-image generation models is the need of a large number of high-quality image-text pairs.
  • 20 Dec 2021 2 repositories listed Syntology ran 9 of 15 samples · 6 unverified
    Diffusion models have recently been shown to generate high-quality synthetic images, especially when paired with a guidance technique to trade off diversity for fidelity.
  • 20 Sep 2023 1 repository listed Syntology ran 3 of 3 samples · 0 unverified · 2 pointer-only (licence)
    This paper presents DreamLLM, a learning framework that first achieves versatile Multimodal Large Language Models (MLLMs) empowered with frequently overlooked synergy between multimodal comprehension and creation.
  • 24 Nov 2022 1 repository listed Syntology ran 7 of 17 samples · 10 unverified
    Unlike the baseline diffusion model used in DALL-E 2, our method seamlessly encodes prior knowledge of the pre-trained CLIP model in its diffusion process by designing a new initialization distribution and a new…
  • 6 Jun 2022 1 repository listed Syntology ran 9 of 13 samples · 4 unverified · 1 pointer-only (licence)
    Our solution leverages a recent text-to-image Latent Diffusion Model (LDM), which speeds up diffusion by operating in a lower-dimensional latent space.
  • 2 Dec 2021 1 repository listed
    We approach text-to-image generation by combining the power of the retrained CLIP representation with an off-the-shelf image generator (GANs), optimizing in the latent space of GAN to find images that achieve maximum…
  • 29 Nov 2021 1 repository listed Syntology ran 14 of 22 samples · 8 unverified · 1 pointer-only (licence)
    Natural language offers a highly intuitive interface for image editing.
  • 22 Nov 2021 1 repository listed
    Unlike other models, BiART can distinguish between image (or text) as a conditional reference and a generation target.

Syntology lines on 9 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections