{"url":"/dataset/heim","name":"HEIM","full_name":"Holistic Evaluation of Text-to-Image Models","description_markdown":"**HEIM** stands for **Holistic Evaluation of Text-To-Image Models**. It is a comprehensive benchmark designed to assess the capabilities and risks of text-to-image generation models. Unlike previous evaluations that primarily focused on image-text alignment and image quality, HEIM considers **12 different aspects** that are crucial for real-world model deployment:\r\n\r\n1. **Image-Text Alignment**\r\n2. **Image Quality**\r\n3. **Aesthetics**\r\n4. **Originality**\r\n5. **Reasoning**\r\n6. **Knowledge**\r\n7. **Bias**\r\n8. **Toxicity**\r\n9. **Fairness**\r\n10. **Robustness**\r\n11. **Multilinguality**\r\n12. **Efficiency**\r\n\r\nBy curating scenarios that encompass these aspects, HEIM evaluates state-of-the-art text-to-image models. Interestingly, no single model excels in all aspects; different models demonstrate strengths in different areas. For transparency, all prompts, generated images, and results are available on the [HEIM website](https://crfm.stanford.edu/helm/heim/latest/) for exploration and study. Additionally, the [GitHub repository](https://github.com/stanford-crfm/helm) provides a collection of models accessible via a unified API, along with metrics beyond accuracy, such as efficiency, bias, and toxicity.","description_withheld":null,"homepage":"https://crfm.stanford.edu/heim/latest","introduced_date":"2023-11-07","introduced_date_note":null,"introduced_by":{"paper":"/paper/holistic-evaluation-of-text-to-image-models-1","title":"Holistic Evaluation of Text-To-Image Models","first_author":"Tony Lee","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["HEIM"],"data_loaders":[],"num_papers_in_archive":13,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}