Papers › SCOT: Self-Supervised Contrastive Pretraining For Zero-Shot Compositional Retrieval

SCOT: Self-Supervised Contrastive Pretraining For Zero-Shot Compositional Retrieval

12 Jan 2025WACV 2025 3arXiv:2501.08347archive 2025-07-28

Bhavin Jawade, Joao V. B. Soares, Kapil Thadani, Deen Dayal Mohan, Amir Erfan Eshratifar, Benjamin Culpepper, Paloma de Juan, Srirangaraj Setlur, Venu Govindaraju

Compositional image retrieval (CIR) is a multimodal learning task where a model combines a query image with a user-provided text modification to retrieve a target image. CIR finds applications in a variety of domains including product retrieval (e-commerce) and web search. Existing methods primarily focus on fully-supervised learning, wherein models are trained on datasets of labeled triplets such as FashionIQ and CIRR. This poses two significant challenges: (i) curating such triplet datasets is labor intensive; and (ii) models lack generalization to unseen objects and domains. In this work, we propose SCOT (Self-supervised COmpositional Training), a novel zero-shot compositional pretraining strategy that combines existing large image-text pair datasets with the generative capabilities of large language models to contrastively train an embedding composition network. Specifically, we show that the text embedding from a large-scale contrastively-pretrained vision-language model can be utilized as proxy target supervision during compositional pretraining, replacing the target image embedding. In zero-shot settings, this strategy surpasses SOTA zero-shot compositional retrieval methods as well as many fully-supervised methods on standard benchmarks such as FashionIQ and CIRR.

PaperPDFConference PDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image RetrievalRetrievalZero-Shot Composed Image Retrieval (ZS-CIR)

1 archive task tag without a task page not shown.

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Zero-Shot Composed Image Retrieval (ZS-CIR) CIRCO SCOT (WACV 2025) mAP@10 37.88 #3 of 43 Archive leaderboard report
Zero-Shot Composed Image Retrieval (ZS-CIR) CIRR SCOT (WACV 2025) R@1 36.82 #8 of 47 Archive leaderboard report
Zero-Shot Composed Image Retrieval (ZS-CIR) CIRR SCOT (WACV 2025) R@10 74.48 #8 of 47 Archive leaderboard report
Zero-Shot Composed Image Retrieval (ZS-CIR) CIRR SCOT (WACV 2025) R@5 64.34 #8 of 47 Archive leaderboard report
Zero-Shot Composed Image Retrieval (ZS-CIR) CIRR SCOT (WACV 2025) R@50 93.42 #8 of 47 Archive leaderboard report
Zero-Shot Composed Image Retrieval (ZS-CIR) Fashion IQ SCOT (WACV 2025) (Recall@10+Recall@50)/2 49.24 #5 of 41 Archive leaderboard report
Zero-Shot Composed Image Retrieval (ZS-CIR) Fashion IQ SCOT (WACV 2025) R@10 38.45 #5 of 41 Archive leaderboard report
Zero-Shot Composed Image Retrieval (ZS-CIR) Fashion IQ SCOT (WACV 2025) R@50 60.03 #5 of 41 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Focus

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections