Papers › PACS: A Dataset for Physical Audiovisual CommonSense Reasoning

PACS: A Dataset for Physical Audiovisual CommonSense Reasoning

21 Mar 2022arXiv:2203.11130archive 2025-07-28

Samuel Yu, Peter Wu, Paul Pu Liang, Ruslan Salakhutdinov, Louis-Philippe Morency

In order for AI to be safely deployed in real-world scenarios such as hospitals, schools, and the workplace, it must be able to robustly reason about the physical world. Fundamental to this reasoning is physical common sense: understanding the physical properties and affordances of available objects, how they can be manipulated, and how they interact with other objects. Physical commonsense reasoning is fundamentally a multi-sensory task, since physical properties are manifested through multiple modalities - two of them being vision and acoustics. Our paper takes a step towards real-world physical commonsense reasoning by contributing PACS: the first audiovisual benchmark annotated for physical commonsense attributes. PACS contains 13,400 question-answer pairs, involving 1,377 unique physical commonsense questions and 1,526 videos. Our dataset provides new opportunities to advance the research field of physical reasoning by bringing audio as a core component of this multimodal problem. Using PACS, we evaluate multiple state-of-the-art models on our new challenging task. While some models show promising results (70% accuracy), they all fall short of human performance (95% accuracy). We conclude the paper by demonstrating the importance of multimodal reasoning and providing possible avenues for future research.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

samuelyu2002/pacs officialmentioned in papermentioned on GitHubjax report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Common Sense ReasoningMultimodal ReasoningPhysical Commonsense Reasoning

Datasets

Introduced by this paper, per the archive.

Physical Audiovisual CommonSense

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Physical Commonsense Reasoning Physical Audiovisual CommonSense Human With Audio (Acc %) 96.3 ± 2.1 #1 of 6 Archive leaderboard report
Physical Commonsense Reasoning Physical Audiovisual CommonSense Human Without Audio (Acc %) 90.5 ± 3.1 #1 of 6 Archive leaderboard report
Physical Commonsense Reasoning Physical Audiovisual CommonSense Merlot Reserve (Large) With Audio (Acc %) 70.1 ± 1.0 #2 of 6 Archive leaderboard report
Physical Commonsense Reasoning Physical Audiovisual CommonSense Merlot Reserve (Large) Without Audio (Acc %) 68.4 ± 0.7 #2 of 6 Archive leaderboard report
Physical Commonsense Reasoning Physical Audiovisual CommonSense CLIP/AudioCLIP With Audio (Acc %) 60.0 ± 0.9 #3 of 6 Archive leaderboard report
Physical Commonsense Reasoning Physical Audiovisual CommonSense CLIP/AudioCLIP Without Audio (Acc %) 56.3 ± 0.7 #3 of 6 Archive leaderboard report
Physical Commonsense Reasoning Physical Audiovisual CommonSense Late Fusion With Audio (Acc %) 55.0 ± 1.1 #4 of 6 Archive leaderboard report
Physical Commonsense Reasoning Physical Audiovisual CommonSense Late Fusion Without Audio (Acc %) 52.5 ± 1.6 #4 of 6 Archive leaderboard report
Physical Commonsense Reasoning Physical Audiovisual CommonSense Majority With Audio (Acc %) 50.4 #5 of 6 Archive leaderboard report
Physical Commonsense Reasoning Physical Audiovisual CommonSense Majority Without Audio (Acc %) 50.4 #5 of 6 Archive leaderboard report
Physical Commonsense Reasoning Physical Audiovisual CommonSense UNITER (Large) Without Audio (Acc %) 60.6 ± 2.2 #6 of 6 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections