Datasets › DrawBench
DrawBench
DrawBench is a comprehensive and challenging benchmark for text-to-image models, introduced by the Imagen research team. Let me provide you with more details:
- Purpose and Context:
- DrawBench serves as an evaluation benchmark specifically designed to assess the performance of text-to-image models.
-
It allows researchers and practitioners to compare different methods and understand their strengths and weaknesses in generating images from textual descriptions.
-
Imagen: Text-to-Image Diffusion Models:
- Imagen is a state-of-the-art text-to-image diffusion model developed by the Google Research Brain Team.
- It combines the power of large transformer language models (such as T5) for understanding text with the strength of diffusion models for high-fidelity image generation.
- Key Discovery: Imagen demonstrates that generic large language models pretrained on text-only corpora are remarkably effective at encoding text for image synthesis.
- Photorealism and Language Understanding: Imagen achieves an unprecedented degree of photorealism and a deep level of language understanding.
- FID Score: It achieves a new state-of-the-art FID (Fréchet Inception Distance) score of 7.27 on the COCO dataset, without ever being trained on COCO.
-
Human Raters' Perception: Human raters find Imagen samples to be on par with the COCO data itself in terms of image-text alignment.
-
DrawBench: A Comprehensive Benchmark:
- DrawBench provides a rigorous evaluation framework for text-to-image models.
- Researchers can compare Imagen with other recent methods, including VQ-GAN+CLIP, Latent Diffusion Models, and DALL-E 2.
-
Human raters prefer Imagen over other models in side-by-side comparisons, considering both sample quality and image-text alignment.
-
Examples from the Imagen Family:
-
Imagen generates diverse and imaginative images based on textual prompts. Here are some examples:
- A strawberry mug filled with white sesame seeds, floating in a dark chocolate sea.
- A brain riding a rocketship heading towards the moon.
- A dragon fruit wearing a karate belt in the snow.
- A small cactus wearing a straw hat and neon sunglasses in the Sahara desert.
- A photo of a Corgi dog riding a bike in Times Square, wearing sunglasses and a beach hat.
- Teddy bears swimming at the Olympics 400m Butterfly event.
- Sprouts in the shape of the text 'Imagen' coming out of a fairytale book.
- A transparent sculpture of a duck made out of glass, in front of a painting of a landscape.
- A single beam of light entering the room from the ceiling, illuminating an easel with a Rembrandt painting of a raccoon.
-
Technical Details:
- Imagen uses a large frozen T5-XXL encoder to encode input text into embeddings.
- The combination of language understanding and diffusion-based image generation results in high-quality, contextually relevant images.
Source: Conversation with Bing, 3/18/2024 (1) Imagen: Text-to-Image Diffusion Models. https://imagen.research.google/. (2) Evaluating Diffusion Models - Hugging Face. https://huggingface.co/docs/diffusers/conceptual/evaluation. (3) shunk031/DrawBench · Datasets at Hugging Face. https://huggingface.co/datasets/shunk031/DrawBench. (4) sayakpaul/drawbench · Datasets at Hugging Face. https://huggingface.co/datasets/sayakpaul/drawbench.
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Text-to-Image Generation | DrawBench | LCM (Curriculum DPO) Aesthetics (Laion Aesthtetics Predictor) 6.1829 | Curriculum Direct Preference Optimization for Diffusion... | croitorualin/curriculum-dpo | 8 | Compare |
Papers archive 2025-07-28
5 shown of 5 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 82. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Curriculum Direct Preference Optimization for Diffusion and Consistency Models | 1 | 2 | 22 May 2024 | ran 3 of 8 samples (5 unverified; 8 pointer-only for licence) |
| Diffusion Model Alignment Using Direct Preference Optimization | 2 | 2 | 21 Nov 2023 | ran 5 of 10 samples (5 unverified; 10 pointer-only for licence) |
| Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference | 5 | 1 | 6 Oct 2023 | ran 1 of 1 samples (0 unverified) |
| Training Diffusion Models with Reinforcement Learning | 3 | 2 | 22 May 2023 | ran 0 of 6 samples (6 unverified) |
| High-Resolution Image Synthesis with Latent Diffusion Models | 41 | 1 | 20 Dec 2021 | ran 19 of 28 samples (9 unverified; 5 pointer-only for licence) |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
No modality tagged.
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- DrawBench
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections