{"url":"/dataset/phybench","name":"PhyBench","full_name":null,"description_markdown":"PhyBench is a comprehensive **Text-to-Image (T2I) evaluation dataset** designed to assess the physical commonsense of T2I models¹. It was introduced by the OpenGVLab and includes **700 prompts** across four primary categories: **mechanics, optics, thermodynamics, and material properties**, covering **31 distinct physical scenarios**¹.\r\n\r\nThe purpose of PhyBench is to evaluate how well T2I models, such as DALL-E 3 and Gemini, can generate images that are consistent with physical principles. The findings from the PhyBench assessments indicate that while these models can often translate text prompts into images, they frequently make errors in depicting physical scenarios correctly, particularly outside of optics¹.\r\n\r\nFor example, when given prompts like \"A cylindrical block of wood placed in front of a mirror\" or \"An apple, a piece of wood, and an iron block in a tank filled with water\", even advanced models like DALL-E 3 and Midjourney have shown to misrepresent the objects or omit them entirely¹.\r\n\r\n(1) GitHub - OpenGVLab/PhyBench. https://github.com/OpenGVLab/PhyBench.\r\n(2) Evaluate your computer's hardware capabilities | Cinebench from Maxon. https://www.maxon.net/en/cinebench.\r\n(3) KegangWangCCNU/PhysBench - GitHub. https://github.com/KegangWangCCNU/PhysBench.","description_withheld":null,"homepage":"https://github.com/OpenGVLab/PhyBench","introduced_date":"2024-06-17","introduced_date_note":null,"introduced_by":{"paper":"/paper/phybench-a-physical-commonsense-benchmark-for","title":"PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models","first_author":"Fanqing Meng","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["PhyBench"],"data_loaders":[],"num_papers_in_archive":6,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}