Browse State-of-the-Art › Text to Image Generation
Text to Image Generation
461 papers with code · 0 benchmarks · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 0 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| DiffusionDB (0 rows) | no rows in the archive | — | — | ||
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 461 papers with code (969 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
28 Nov 2017 20 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedIn this paper, we propose an Attentional Generative Adversarial Network (AttnGAN) that allows attention-driven, multi-stage refinement for fine-grained text-to-image generation.
-
10 Feb 2023 12 repositories listed Syntology ran 2 of 4 samples · 2 unverifiedControlNet locks the production-ready large diffusion models, and reuses their deep and robust encoding layers pretrained with billions of images as a strong backbone to learn a diverse set of conditional controls.
-
24 Feb 2021 12 repositories listed Syntology ran 7 of 7 samples · 0 unverified · 3 pointer-only (licence)Text-to-image generation has traditionally focused on finding better modeling assumptions for training on a fixed dataset.
-
2 Aug 2022 9 repositories listed Syntology ran 10 of 13 samples · 3 unverified · 1 pointer-only (licence)Yet, it is unclear how such freedom can be exercised to generate images of specific unique concepts, modify their appearance, or compose them in new roles and novel scenes.
-
20 Feb 2023 6 repositories listedRecent large-scale generative models learned on big data are capable of synthesizing incredible images yet suffer from limited controllability.
-
17 Nov 2022 6 repositories listed Syntology ran 15 of 20 samples · 5 unverifiedWe propose a method for editing images from human instructions: given an input image and a written instruction that tells the model what to do, our model follows these instructions to edit the image.
-
6 Oct 2023 5 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedInspired by Consistency Models (song et al.), we propose Latent Consistency Models (LCMs), enabling swift inference with minimal steps on any pre-trained LDMs, including Stable Diffusion (rombach et al).
-
2 Jan 2023 5 repositories listed Syntology ran 18 of 21 samples · 3 unverified · 9 pointer-only (licence)Compared to pixel-space diffusion models, such as Imagen and DALL-E 2, Muse is significantly more efficient due to the use of discrete tokens and requiring fewer sampling iterations; compared to autoregressive models,…
-
1 Jun 2023 4 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedPre-trained large text-to-image models synthesize impressive images with an appropriate use of text prompts.
-
17 Apr 2023 4 repositories listed Syntology ran 4 of 7 samples · 3 unverified · 1 pointer-only (licence)Despite the success in large-scale text-to-image generation and text-conditioned image editing, existing methods still struggle to produce consistent generation and editing results.
-
14 Nov 2022 4 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedRecent advancements in the domain of text-to-image synthesis have culminated in a multitude of enhancements pertaining to quality, fidelity, and diversity.
-
26 May 2021 4 repositories listed Syntology ran 4 of 6 samples · 2 unverifiedText-to-Image generation in the general domain has long been an open problem, which requires both a powerful generative model and cross-modal understanding.
-
5 Dec 2023 3 repositories listedWe demonstrate the usage of state-of-the-art text-to-image architectures in the context of laparoscopic imaging with regard to the surgical removal of the gallbladder as an example.
-
12 Apr 2023 3 repositories listed Syntology ran 3 of 12 samples · 9 unverified · 2 pointer-only (licence)We present a comprehensive solution to learn and improve text-to-image models from human preference feedback.
-
31 Mar 2023 3 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 1 pointer-only (licence)Recent breakthroughs in the field of language-guided image generation have yielded impressive achievements, enabling the creation of high-quality and diverse images based on user instructions.
-
12 Mar 2023 3 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Inspired by the unified view, UniDiffuser learns all distributions simultaneously with a minimal modification to the original diffusion model -- perturbs data in all modalities instead of a single modality, inputs…
-
22 Feb 2023 3 repositories listed Syntology ran 4 of 24 samples · 20 unverifiedIn this work, we build upon these ideas using the score-based interpretation of diffusion models, and explore alternative ways to condition, modify, and reuse diffusion models for tasks involving compositional…
-
16 Feb 2023 3 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)In this work, we present MultiDiffusion, a unified framework that enables versatile and controllable image generation, using a pre-trained text-to-image diffusion model, without any further training or finetuning.
-
19 Dec 2022 3 repositories listed Syntology ran 1 of 9 samples · 8 unverifiedInstead of laborious human engineering, we propose prompt adaptation, a general framework that automatically adapts original user input to model-preferred prompts.
-
2 Nov 2022 3 repositories listed Syntology ran 5 of 6 samples · 1 unverified · 6 pointer-only (licence)The commonly-used fast sampler for guided sampling is DDIM, a first-order diffusion ODE solver that generally needs 100 to 250 steps for high-quality samples.
-
25 Sep 2022 3 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)We evaluate U-ViT in unconditional and class-conditional image generation, as well as text-to-image generation tasks, where U-ViT is comparable if not superior to a CNN-based U-Net of a similar size.
-
27 Nov 2021 3 repositories listed Syntology ran 4 of 18 samples · 14 unverifiedOne of the major challenges in training text-to-image generation models is the need of a large number of high-quality image-text pairs.
-
11 May 2020 3 repositories listedThis can be done by conditioning the model on additional information.
-
24 Nov 2018 3 repositories listedConditional text-to-image generation is an active area of research, with many possible applications.
-
3 Jun 2025 2 repositories listedAlthough existing unified models achieve strong performance in vision-language understanding and text-to-image generation, they remain limited in addressing image perception and manipulation -- capabilities increasingly…
-
28 May 2025 2 repositories listed Syntology ran 2 of 3 samples · 1 unverifiedThen, a single-stream sparse DiT structure with dynamic MoE architecture is adopted to trigger multi-model interaction for image generation in a cost-efficient manner.
-
26 May 2025 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)To ensure the data quality, we employ a multi-stage pipeline that integrates a cutting-edge vision-language model, a detection model, a segmentation model, alongside task-specific in-painting procedures and strict…
-
1 May 2025 2 repositories listed Syntology ran 2 of 4 samples · 2 unverified · 3 pointer-only (licence)By applying our reasoning strategies to the baseline model, Janus-Pro, we achieve superior performance with 13% improvement on T2I-CompBench and 19% improvement on the WISE benchmark, even surpassing the…
-
14 Apr 2025 2 repositories listedIn general cases, existing text-to-image generation models excel in producing high-quality images; however, they struggle to capture diverse characteristics and faithful details of specific domains, particularly Chinese…
-
21 Mar 2025 2 repositories listed Syntology ran 4 of 5 samples · 1 unverifiedWe then propose a new sampling strategy based on our Halton scheduler instead of the original Confidence scheduler.
Syntology lines on 24 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections