{"url":"/dataset/phy-q","name":"Phy-Q","full_name":null,"description_markdown":"**Phy-Q** is a benchmark that requires an agent to reason about physical scenarios and take an action accordingly. Inspired by the physical knowledge acquired in infancy and the capabilities required for robots to operate in real-world environments, the authors identify 15 essential physical scenarios. For each scenario, a wide variety of distinct task templates are created, and all the task templates within the same scenario can be solved by using one specific physical rule. \r\n\r\nBy having such a design, two distinct levels of generalization can be evaluated, namely the local generalization and the broad generalization.  The benchmark gives a Phy-Q (physical reasoning quotient) score that reflects the physical reasoning ability of the agents.","description_withheld":null,"homepage":"https://github.com/phy-q/benchmark","introduced_date":"2021-08-31","introduced_date_note":null,"introduced_by":{"paper":"/paper/phy-q-a-benchmark-for-physical-reasoning","title":"Phy-Q as a measure for physical reasoning intelligence","first_author":"Cheng Xue","url":null},"license":null,"modalities":[{"name":"Environment","url":"/datasets/modality/environment"}],"tasks":[],"languages":[],"variants":["Phy-Q"],"data_loaders":[],"num_papers_in_archive":4,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}