{"url":"/dataset/genai-bench","name":"GenAI-Bench","full_name":"GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual Generation","description_markdown":"GenAI-Bench is a benchmarking framework designed to evaluate and improve compositional text-to-visual generation models. It was developed by researchers from Carnegie Mellon University and Meta. The key aspects of GenAI-Bench include:\r\n\r\n- **Compositional Text-to-Visual Generation**: It focuses on the ability of generative models to handle compositional text prompts that involve attributes, relationships, and higher-order reasoning such as logic and comparison¹.\r\n  \r\n- **Diverse Text Prompts**: GenAI-Bench uses 1,600 text prompts collected from professional graphic designers to cover a wide range of compositional reasoning skills¹.\r\n  \r\n- **Human Studies**: The framework includes human annotators who rate the performance of leading generative models like DALL-E 3, Stable Diffusion, and others based on image-text or video-text alignment¹.\r\n  \r\n- **Automated Evaluation Metrics**: It aims to benchmark automated evaluation metrics that measure the alignment between an image and a text prompt².\r\n  \r\n- **Improving Generation**: GenAI-Bench also explores how VQAScore, an automated metric, can improve image generation by selecting the highest-scoring images from generated candidates¹.\r\n\r\nOverall, GenAI-Bench provides a comprehensive and challenging testbed for state-of-the-art text-to-visual generative models, pushing the boundaries of what these models can achieve in terms of understanding and creating complex visual compositions based on textual descriptions.\r\n\r\n(1) GenAI-Bench: Evaluating and Improving Compositional Text-to-Visual .... https://linzhiqiu.github.io/papers/genai_bench/.\r\n(2) GenAI-Bench: A Holistic Benchmark for Compositional Text-to-Visual .... https://openreview.net/pdf?id=hJm7qnW3ym.\r\n(3) Evaluating Text-to-Visual Generation with Image-to-Text Generation. https://linzhiqiu.github.io/papers/vqascore/.","description_withheld":null,"homepage":"https://huggingface.co/datasets/BaiqiL/GenAI-Bench","introduced_date":"2024-06-19","introduced_date_note":null,"introduced_by":null,"license":{"name":"apache-2.0","url":"https://www.apache.org/licenses/LICENSE-2.0"},"modalities":[],"tasks":[],"languages":[],"variants":["GenAI-Bench"],"data_loaders":[],"num_papers_in_archive":13,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}