Datasets › GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

19 Jun 2024 archive 2025-07-28

GenAI-Bench benchmark consists of 1,600 challenging real-world text prompts sourced from professional designers. Compared to benchmarks such as PartiPrompt and T2I-CompBench, GenAI-Bench captures a wider range of aspects in the compositional text-to-visual generation, ranging from basic (scene, attribute, relation) to advanced (counting, comparison, differentiation, logic). GenAI-Bench benchmark also collects human alignment ratings (1-to-5 Likert scales) on images and videos generated by ten leading models, such as Stable Diffusion, DALL-E 3, Midjourney v6, Pika v1, and Gen2.

GenAI-Bench:

  1. Prompt: 1600 prompts sourced from professional designers.
  2. Compositional Skill Tags: Multiple compositional tags for each prompt. The compositional skill tags are categorized into Basic Skill and Advanced Skill. For detailed definitions and examples, please refer to our paper.
  3. Images: Generated images are collected from DALLE_3, DeepFloyd_I_XL_v1, Midjourney_6, SDXL_2_1, SDXL_Base and SDXL_Turbo.
  4. Human Ratings: 1-to-5 Likert scale ratings for each image.

Benchmarks archive 2025-07-28

No leaderboard in the archive resolves to this dataset.

Papers archive 2025-07-28

No paper in the archive has a leaderboard row on this dataset.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

No task tagged in the archive.

License archive 2025-07-28

apache-2.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections