Papers › ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model...

ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective Reasoning

13 Mar 2025arXiv:2503.10166archive 2025-07-28

Pengfei Luo, Jingbo Zhou, Tong Xu, Yuan Xia, Linli Xu, Enhong Chen

With the proliferation of images in online content, language-guided image retrieval (LGIR) has emerged as a research hotspot over the past decade, encompassing a variety of subtasks with diverse input forms. While the development of large multimodal models (LMMs) has significantly facilitated these tasks, existing approaches often address them in isolation, requiring the construction of separate systems for each task. This not only increases system complexity and maintenance costs, but also exacerbates challenges stemming from language ambiguity and complex image content, making it difficult for retrieval systems to provide accurate and reliable results. To this end, we propose ImageScope, a training-free, three-stage framework that leverages collective reasoning to unify LGIR tasks. The key insight behind the unification lies in the compositional nature of language, which transforms diverse LGIR tasks into a generalized text-to-image retrieval process, along with the reasoning of LMMs serving as a universal verification to refine the results. To be specific, in the first stage, we improve the robustness of the framework by synthesizing search intents across varying levels of semantic granularity using chain-of-thought (CoT) reasoning. In the second and third stages, we then reflect on retrieval results by verifying predicate propositions locally, and performing pairwise evaluations globally. Experiments conducted on six LGIR datasets demonstrate that ImageScope outperforms competitive baselines. Comprehensive evaluations and ablation studies further confirm the effectiveness of our design.

PaperPDFConference PDFCode

Code

pengfei-luo/ImageScope officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image RetrievalRetrievalZero-Shot Composed Image Retrieval (ZS-CIR)

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Chat-based Image Retrieval VisDial ImageScope (CLIP-ViT-L/14) Hits@10 on 10 Round 79.89 #3 of 3 Archive leaderboard report
Zero-Shot Composed Image Retrieval (ZS-CIR) CIRCO ImageScope (CLIP-ViT-L/14) MAP@5 28.36 #14 of 43 Archive leaderboard report
Zero-Shot Composed Image Retrieval (ZS-CIR) CIRCO ImageScope (CLIP-ViT-L/14) mAP@10 29.23 #14 of 43 Archive leaderboard report
Zero-Shot Composed Image Retrieval (ZS-CIR) CIRCO ImageScope (CLIP-ViT-L/14) mAP@25 30.81 #14 of 43 Archive leaderboard report
Zero-Shot Composed Image Retrieval (ZS-CIR) CIRCO ImageScope (CLIP-ViT-L/14) mAP@50 31.88 #14 of 43 Archive leaderboard report
Zero-Shot Composed Image Retrieval (ZS-CIR) CIRR ImageScope (CLIP-ViT-L/14) R@1 39.37 #3 of 47 Archive leaderboard report
Zero-Shot Composed Image Retrieval (ZS-CIR) CIRR ImageScope (CLIP-ViT-L/14) R@10 78.05 #3 of 47 Archive leaderboard report
Zero-Shot Composed Image Retrieval (ZS-CIR) CIRR ImageScope (CLIP-ViT-L/14) R@5 67.54 #3 of 47 Archive leaderboard report
Zero-Shot Composed Image Retrieval (ZS-CIR) CIRR ImageScope (CLIP-ViT-L/14) R@50 92.94 #3 of 47 Archive leaderboard report
Zero-Shot Composed Image Retrieval (ZS-CIR) Fashion IQ ImageScope (CLIP-ViT-L/14) R@10 31.36 #41 of 41 Archive leaderboard report
Zero-Shot Composed Image Retrieval (ZS-CIR) Fashion IQ ImageScope (CLIP-ViT-L/14) R@50 50.78 #41 of 41 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections