{"url":"/dataset/genai-bench-evaluating-and-improving","name":"GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation","full_name":"GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation","description_markdown":"GenAI-Bench benchmark consists of 1,600 challenging real-world text prompts sourced from professional designers. Compared to benchmarks such as PartiPrompt and T2I-CompBench, GenAI-Bench captures a wider range of aspects in the compositional text-to-visual generation, ranging from basic (scene, attribute, relation) to advanced (counting, comparison, differentiation, logic). GenAI-Bench benchmark also collects human alignment ratings (1-to-5 Likert scales) on images and videos generated by ten leading models, such as Stable Diffusion, DALL-E 3, Midjourney v6, Pika v1, and Gen2.\r\n\r\nGenAI-Bench:\r\n\r\n1. Prompt: 1600 prompts sourced from professional designers.\r\n2. Compositional Skill Tags: Multiple compositional tags for each prompt. The compositional skill tags are categorized into Basic Skill and Advanced Skill. For detailed definitions and examples, please refer to our paper.\r\n3. Images: Generated images are collected from DALLE_3, DeepFloyd_I_XL_v1, Midjourney_6, SDXL_2_1, SDXL_Base and SDXL_Turbo.\r\n4. Human Ratings: 1-to-5 Likert scale ratings for each image.","description_withheld":null,"homepage":"https://linzhiqiu.github.io/papers/genai_bench/","introduced_date":"2024-06-19","introduced_date_note":null,"introduced_by":null,"license":{"name":"apache-2.0","url":"https://www.apache.org/licenses/LICENSE-2.0"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Videos","url":"/datasets/modality/videos"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}