Papers › Think before You Simulate: Symbolic Reasoning to Orchestrate Neural Computation for...
Think before You Simulate: Symbolic Reasoning to Orchestrate Neural Computation for Counterfactual Question Answering
Adam Ishay, Zhun Yang, Joohyung Lee, Ilgu Kang, Dongjae Lim
Causal and temporal reasoning about video dynamics is a challenging problem. While neuro-symbolic models that combine symbolic reasoning with neural-based perception and prediction have shown promise, they exhibit limitations, especially in answering counterfactual questions. This paper introduces a method to enhance a neuro-symbolic model for counterfactual reasoning, leveraging symbolic reasoning about causal relations among events. We define the notion of a causal graph to represent such relations and use Answer Set Programming (ASP), a declarative logic programming method, to find how to coordinate perception and simulation modules. We validate the effectiveness of our approach on two benchmarks, CLEVRER and CRAFT. Our enhancement achieves state-of-the-art performance on the CLEVRER challenge, significantly outperforming existing models. In the case of the CRAFT benchmark, we leverage a large pre-trained language model, such as GPT-3.5 and GPT-4, as a proxy for a dynamics simulator. Our findings show that this method can further improve its performance on counterfactual questions by providing alternative prompts instructed by symbolic causal reasoning.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
1 archive task tag without a task page not shown.
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Visual Reasoning | CLEVRER | AI Core | Average-per ques. | 95.24 | #1 of 13 | Archive leaderboard | report |
| Visual Reasoning | CLEVRER | AI Core | Counterfactual-per opt. | 96.61 | #1 of 13 | Archive leaderboard | report |
| Visual Reasoning | CLEVRER | AI Core | Counterfactual-per ques. | 90.72 | #1 of 13 | Archive leaderboard | report |
| Visual Reasoning | CLEVRER | AI Core | Descriptive | 96.46 | #1 of 13 | Archive leaderboard | report |
| Visual Reasoning | CLEVRER | AI Core | Explanatory-per opt. | 99.94 | #1 of 13 | Archive leaderboard | report |
| Visual Reasoning | CLEVRER | AI Core | Explanatory-per ques. | 99.81 | #1 of 13 | Archive leaderboard | report |
| Visual Reasoning | CLEVRER | AI Core | Predictive-per opt. | 93.96 | #1 of 13 | Archive leaderboard | report |
| Visual Reasoning | CLEVRER | AI Core | Predictive-per ques. | 93.96 | #1 of 13 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections