Datasets › DrawBench

DrawBench

Introduced by Chitwan Saharia et al. in Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding23 May 2022 archive 2025-07-28

DrawBench is a comprehensive and challenging benchmark for text-to-image models, introduced by the Imagen research team. Let me provide you with more details:

  1. Purpose and Context:
  2. DrawBench serves as an evaluation benchmark specifically designed to assess the performance of text-to-image models.
  3. It allows researchers and practitioners to compare different methods and understand their strengths and weaknesses in generating images from textual descriptions.

  4. Imagen: Text-to-Image Diffusion Models:

  5. Imagen is a state-of-the-art text-to-image diffusion model developed by the Google Research Brain Team.
  6. It combines the power of large transformer language models (such as T5) for understanding text with the strength of diffusion models for high-fidelity image generation.
  7. Key Discovery: Imagen demonstrates that generic large language models pretrained on text-only corpora are remarkably effective at encoding text for image synthesis.
  8. Photorealism and Language Understanding: Imagen achieves an unprecedented degree of photorealism and a deep level of language understanding.
  9. FID Score: It achieves a new state-of-the-art FID (Fréchet Inception Distance) score of 7.27 on the COCO dataset, without ever being trained on COCO.
  10. Human Raters' Perception: Human raters find Imagen samples to be on par with the COCO data itself in terms of image-text alignment.

  11. DrawBench: A Comprehensive Benchmark:

  12. DrawBench provides a rigorous evaluation framework for text-to-image models.
  13. Researchers can compare Imagen with other recent methods, including VQ-GAN+CLIP, Latent Diffusion Models, and DALL-E 2.
  14. Human raters prefer Imagen over other models in side-by-side comparisons, considering both sample quality and image-text alignment.

  15. Examples from the Imagen Family:

  16. Imagen generates diverse and imaginative images based on textual prompts. Here are some examples:

    • A strawberry mug filled with white sesame seeds, floating in a dark chocolate sea.
    • A brain riding a rocketship heading towards the moon.
    • A dragon fruit wearing a karate belt in the snow.
    • A small cactus wearing a straw hat and neon sunglasses in the Sahara desert.
    • A photo of a Corgi dog riding a bike in Times Square, wearing sunglasses and a beach hat.
    • Teddy bears swimming at the Olympics 400m Butterfly event.
    • Sprouts in the shape of the text 'Imagen' coming out of a fairytale book.
    • A transparent sculpture of a duck made out of glass, in front of a painting of a landscape.
    • A single beam of light entering the room from the ceiling, illuminating an easel with a Rembrandt painting of a raccoon.
  17. Technical Details:

  18. Imagen uses a large frozen T5-XXL encoder to encode input text into embeddings.
  19. The combination of language understanding and diffusion-based image generation results in high-quality, contextually relevant images.

Source: Conversation with Bing, 3/18/2024 (1) Imagen: Text-to-Image Diffusion Models. https://imagen.research.google/. (2) Evaluating Diffusion Models - Hugging Face. https://huggingface.co/docs/diffusers/conceptual/evaluation. (3) shunk031/DrawBench · Datasets at Hugging Face. https://huggingface.co/datasets/shunk031/DrawBench. (4) sayakpaul/drawbench · Datasets at Hugging Face. https://huggingface.co/datasets/sayakpaul/drawbench.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Text-to-Image Generation DrawBench LCM (Curriculum DPO) Aesthetics (Laion Aesthtetics Predictor) 6.1829 Curriculum Direct Preference Optimization for Diffusion... croitorualin/curriculum-dpo 8 Compare

Papers archive 2025-07-28

5 shown of 5 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 82. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Curriculum Direct Preference Optimization for Diffusion and Consistency Models 1 2 22 May 2024 ran 3 of 8 samples (5 unverified; 8 pointer-only for licence)
Diffusion Model Alignment Using Direct Preference Optimization 2 2 21 Nov 2023 ran 5 of 10 samples (5 unverified; 10 pointer-only for licence)
Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference 5 1 6 Oct 2023 ran 1 of 1 samples (0 unverified)
Training Diffusion Models with Reinforcement Learning 3 2 22 May 2023 ran 0 of 6 samples (6 unverified)
High-Resolution Image Synthesis with Latent Diffusion Models 41 1 20 Dec 2021 ran 19 of 28 samples (9 unverified; 5 pointer-only for licence)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • DrawBench

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections