{"url":"/dataset/drawbench","name":"DrawBench","full_name":null,"description_markdown":"**DrawBench** is a comprehensive and challenging benchmark for text-to-image models, introduced by the **Imagen** research team. Let me provide you with more details:\r\n\r\n1. **Purpose and Context**:\r\n   - **DrawBench** serves as an evaluation benchmark specifically designed to assess the performance of text-to-image models.\r\n   - It allows researchers and practitioners to compare different methods and understand their strengths and weaknesses in generating images from textual descriptions.\r\n\r\n2. **Imagen: Text-to-Image Diffusion Models**:\r\n   - **Imagen** is a state-of-the-art text-to-image diffusion model developed by the Google Research Brain Team.\r\n   - It combines the power of large transformer language models (such as T5) for understanding text with the strength of diffusion models for high-fidelity image generation.\r\n   - **Key Discovery**: Imagen demonstrates that generic large language models pretrained on text-only corpora are remarkably effective at encoding text for image synthesis.\r\n   - **Photorealism and Language Understanding**: Imagen achieves an unprecedented degree of photorealism and a deep level of language understanding.\r\n   - **FID Score**: It achieves a new state-of-the-art FID (Fréchet Inception Distance) score of **7.27** on the COCO dataset, without ever being trained on COCO.\r\n   - **Human Raters' Perception**: Human raters find Imagen samples to be on par with the COCO data itself in terms of image-text alignment.\r\n\r\n3. **DrawBench: A Comprehensive Benchmark**:\r\n   - **DrawBench** provides a rigorous evaluation framework for text-to-image models.\r\n   - Researchers can compare Imagen with other recent methods, including **VQ-GAN+CLIP**, **Latent Diffusion Models**, and **DALL-E 2**.\r\n   - Human raters prefer Imagen over other models in side-by-side comparisons, considering both sample quality and image-text alignment.\r\n\r\n4. **Examples from the Imagen Family**:\r\n   - Imagen generates diverse and imaginative images based on textual prompts. Here are some examples:\r\n     - A strawberry mug filled with white sesame seeds, floating in a dark chocolate sea.\r\n     - A brain riding a rocketship heading towards the moon.\r\n     - A dragon fruit wearing a karate belt in the snow.\r\n     - A small cactus wearing a straw hat and neon sunglasses in the Sahara desert.\r\n     - A photo of a Corgi dog riding a bike in Times Square, wearing sunglasses and a beach hat.\r\n     - Teddy bears swimming at the Olympics 400m Butterfly event.\r\n     - Sprouts in the shape of the text 'Imagen' coming out of a fairytale book.\r\n     - A transparent sculpture of a duck made out of glass, in front of a painting of a landscape.\r\n     - A single beam of light entering the room from the ceiling, illuminating an easel with a Rembrandt painting of a raccoon.\r\n\r\n5. **Technical Details**:\r\n   - Imagen uses a large frozen **T5-XXL** encoder to encode input text into embeddings.\r\n   - The combination of language understanding and diffusion-based image generation results in high-quality, contextually relevant images.\r\n\r\nSource: Conversation with Bing, 3/18/2024\r\n(1) Imagen: Text-to-Image Diffusion Models. https://imagen.research.google/.\r\n(2) Evaluating Diffusion Models - Hugging Face. https://huggingface.co/docs/diffusers/conceptual/evaluation.\r\n(3) shunk031/DrawBench · Datasets at Hugging Face. https://huggingface.co/datasets/shunk031/DrawBench.\r\n(4) sayakpaul/drawbench · Datasets at Hugging Face. https://huggingface.co/datasets/sayakpaul/drawbench.","description_withheld":null,"homepage":"https://huggingface.co/datasets/shunk031/DrawBench","introduced_date":"2022-05-23","introduced_date_note":null,"introduced_by":{"paper":"/paper/photorealistic-text-to-image-diffusion-models","title":"Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding","first_author":"Chitwan Saharia","url":null},"license":null,"modalities":[],"tasks":[{"name":"Text-to-Image Generation","url":"/task/text-to-image-generation","datasets_with_task":"/datasets/task/text-to-image-generation"}],"languages":[],"variants":["DrawBench"],"data_loaders":[],"num_papers_in_archive":82,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/text-to-image-generation-on-drawbench","task":"Text-to-Image Generation","dataset_variant":"DrawBench","rows":8,"metrics":["Aesthetics (Laion Aesthtetics Predictor)","Human Preference Alignement (HPSv2)","Text Alignement (SentenceBERT)"],"first_row_in_archive_order":{"model":"LCM (Curriculum DPO)","paper":"/paper/curriculum-direct-preference-optimization-for","metrics":{"Aesthetics (Laion Aesthtetics Predictor)":"6.1829","Human Preference Alignement (HPSv2)":"0.2851","Text Alignement (SentenceBERT)":"0.5812"},"code_links":[{"title":"croitorualin/curriculum-dpo","url":"https://github.com/croitorualin/curriculum-dpo"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/curriculum-direct-preference-optimization-for","title":"Curriculum Direct Preference Optimization for Diffusion and Consistency Models","date":"2024-05-22","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":8,"samples_ran":3,"samples_unverified":5,"pointer_only_for_licence":8,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/diffusion-model-alignment-using-direct","title":"Diffusion Model Alignment Using Direct Preference Optimization","date":"2023-11-21","rows_on_this_dataset":2,"code_links":2,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":10,"samples_ran":5,"samples_unverified":5,"pointer_only_for_licence":10,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/latent-consistency-models-synthesizing-high","title":"Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference","date":"2023-10-06","rows_on_this_dataset":1,"code_links":5,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/training-diffusion-models-with-reinforcement","title":"Training Diffusion Models with Reinforcement Learning","date":"2023-05-22","rows_on_this_dataset":2,"code_links":3,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":6,"samples_ran":4,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/high-resolution-image-synthesis-with-latent","title":"High-Resolution Image Synthesis with Latent Diffusion Models","date":"2021-12-20","rows_on_this_dataset":1,"code_links":41,"syntology":{"read_at":"2026-09-25T09:33:49+00:00","samples_harvested":28,"samples_ran":21,"samples_unverified":7,"pointer_only_for_licence":5,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":5,"samples_harvested":53,"samples_ran":34,"samples_unverified":19,"pointer_only_for_licence":23,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}