Papers › Compositional Chain-of-Thought Prompting for Large Multimodal Models

Compositional Chain-of-Thought Prompting for Large Multimodal Models

27 Nov 2023CVPR 2024 1arXiv:2311.17076archive 2025-07-28

Chancharik Mitra, Brandon Huang, Trevor Darrell, Roei Herzig

The combination of strong visual backbones and Large Language Model (LLM) reasoning has led to Large Multimodal Models (LMMs) becoming the current standard for a wide range of vision and language (VL) tasks. However, recent research has shown that even the most advanced LMMs still struggle to capture aspects of compositional visual reasoning, such as attributes and relationships between objects. One solution is to utilize scene graphs (SGs)--a formalization of objects and their relations and attributes that has been extensively used as a bridge between the visual and textual domains. Yet, scene graph data requires scene graph annotations, which are expensive to collect and thus not easily scalable. Moreover, finetuning an LMM based on SG data can lead to catastrophic forgetting of the pretraining objective. To overcome this, inspired by chain-of-thought methods, we propose Compositional Chain-of-Thought (CCoT), a novel zero-shot Chain-of-Thought prompting method that utilizes SG representations in order to extract compositional knowledge from an LMM. Specifically, we first generate an SG using the LMM, and then use that SG in the prompt to produce a response. Through extensive experiments, we find that the proposed CCoT approach not only improves LMM performance on several vision and language VL compositional benchmarks but also improves the performance of several popular LMMs on general multimodal benchmarks, without the need for fine-tuning or annotated ground-truth SGs. Code: https://github.com/chancharikmitra/CCoT

PaperPDFConference PDFCodeCode Syntology ran

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

For agents, Syntology's MCP tool lists every function and class Syntology harvested from this paper and whether it ran (how to connect): get_harvested_code_for_paper(arxiv_id="2311.17076")

Code

Syntology Ran 3 of 4 code samples harvested from 1 repository linked to this paper; 1 has no recorded run. Of those that ran: 1 ran · violated contract; 2 ran · our draft was wrong.

By repository: official repository: 4 samples from 1 repository, 3 ran. The run record, sample by sample. “Ran” means executed on a synthesized input, not that the code is correct or reproduces the paper.

chancharikmitra/ccot officialmentioned in papermentioned on GitHubpytorchMIT report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

4 samples harvested; 3 ran; 0 honoured the contract we drafted; 1 has no recorded run. Read from Syntology's graph 2026-09-24; that is when this build read the record, not when the samples ran.

1ran · violated contract
2ran · our draft was wrong
1unverified

Licence: 0 of the 4 samples are pointer only, meaning Syntology does not serve that copy's text. This page shows no code text for any sample; each one links to its file in the repository.

Harvested from chancharikmitra/ccot. “Ran” means the sample executed on a synthesized input. It does not mean the output is correct, and nothing here reproduces the paper's results. “Honoured” and “violated” refer to a contract Syntology drafted from the code itself; “our draft was wrong” and “fixture could not drive it” are failures of Syntology's instrument, not of the code.

Each sample ends with its code_sha256, Syntology's identity for that exact code. An agent fetches the stored sample with Syntology's MCP tool get_code(code_sha256="…") (how to connect); click an identity to copy that call.

Repository labels, per sample. official repository: The archive marks this repository official for the paper. named in the paper: The archive records that the paper mentions this repository; it is not marked official. community (archive-listed): In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper. found in paper text by Syntology: Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted. community: Not in the archive's code links for this paper; a community repository Syntology harvested. Samples from a repository marked official are listed first. Licence labels name the repository's licence as recorded at harvest. “Pointer only” means Syntology does not serve that copy's text, for one of four reasons: no licence file was found; the licence was not identified; the licence is recorded as permissive but that copy's record is not marked cleared; or the licence is outside the permissive list Syntology serves text under (MIT, Apache-2.0, BSD and similar). Some licences outside that list permit redistribution, such as WTFPL, and GPL-3.0 under its conditions; they are simply not on the list. Hover a licence label for the reason. File links open the file on GitHub at the default branch, which may have changed since the harvest.

get_chunk chancharikmitra/ccot/GPT-4V/Sphinx_bench.py official repository ran · our draft was wrong fingerprinted MIT (permissive) · 42a46570620cd9fa · report
is_none chancharikmitra/ccot/InstructBLIP-13b/InstructBLIP_MMBench.py official repository ran · violated contract MIT (permissive) · bae18947b56f2be1 · report
split_list chancharikmitra/ccot/GPT-4V/Sphinx_bench.py official repository ran · our draft was wrong fingerprinted MIT (permissive) · 076c252c52cbb161 · report
get_ans chancharikmitra/ccot/GPT-4V/GPT4V_Whoops.py official repository unverified MIT (permissive) · b2a7f7ca0f25da4c · report

Tasks

Language ModellingLarge Language ModelVisual Reasoning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Reasoning Winoground LLaVA-1.5-CCoT Group Score 22.3 #31 of 114 Archive leaderboard report
Visual Reasoning Winoground LLaVA-1.5-CCoT Image Score 35.5 #31 of 114 Archive leaderboard report
Visual Reasoning Winoground LLaVA-1.5-CCoT Text Score 42.0 #31 of 114 Archive leaderboard report
Visual Reasoning Winoground LLaVA-1.5 Group Score 20.1 #49 of 114 Archive leaderboard report
Visual Reasoning Winoground LLaVA-1.5 Image Score 33.3 #49 of 114 Archive leaderboard report
Visual Reasoning Winoground LLaVA-1.5 Text Score 36.0 #49 of 114 Archive leaderboard report
Visual Reasoning Winoground LLaVA-1.5-ZS-CoT Group Score 12.3 #79 of 114 Archive leaderboard report
Visual Reasoning Winoground LLaVA-1.5-ZS-CoT Image Score 22.5 #79 of 114 Archive leaderboard report
Visual Reasoning Winoground LLaVA-1.5-ZS-CoT Text Score 28.0 #79 of 114 Archive leaderboard report
Visual Reasoning Winoground InstructBLIP-CCoT Group Score 8.3 #97 of 114 Archive leaderboard report
Visual Reasoning Winoground InstructBLIP-CCoT Image Score 21.3 #97 of 114 Archive leaderboard report
Visual Reasoning Winoground InstructBLIP-CCoT Text Score 21.0 #97 of 114 Archive leaderboard report
Visual Reasoning Winoground InstructBLIP-ZS-CoT Group Score 4.0 #113 of 114 Archive leaderboard report
Visual Reasoning Winoground InstructBLIP-ZS-CoT Image Score 16.3 #113 of 114 Archive leaderboard report
Visual Reasoning Winoground InstructBLIP-ZS-CoT Text Score 9.3 #113 of 114 Archive leaderboard report
Visual Reasoning Winoground InstructBLIP Group Score 3.3 #114 of 114 Archive leaderboard report
Visual Reasoning Winoground InstructBLIP Image Score 11.5 #114 of 114 Archive leaderboard report
Visual Reasoning Winoground InstructBLIP Text Score 7.0 #114 of 114 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections