Datasets › CIRR

CIRR (Compose Image Retrieval on Real-life images)

Introduced by Zheyuan Liu et al. in Image Retrieval on Real-life Images with Pre-trained Vision-and-Language Models9 Aug 2021 archive 2025-07-28

Composed Image Retrieval (or, Image Retreival conditioned on Language Feedback) is a relatively new retrieval task, where an input query consists of an image and short textual description of how to modify the image.

For humans, the advantage of a bi-modal query is clear: some concepts and attributes are more succinctly described visually, others through language. By cross-referencing the two modalities, a reference image can capture the general gist of a scene, while the text can specify finer details.

We identify a major challenge of this task as the inherent ambiguity in knowing what information is important (typically one object of interest in the scene) and what can be ignored (e.g., the background and other irrelevant objects).

We release the first dataset of open-domain, real-life images with human-generated modification sentences, which support research on one-shot composed image retrieval, dialogue systems, fine-grained visiolinguistic reasoning, and more.

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

30 shown of 35 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 61. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
TMCIR: Token Merge Benefits Composed Image Retrieval 0 1 15 Apr 2025 not harvested
CoLLM: A Large Language Model for Composed Image Retrieval 1 3 25 Mar 2025 not harvested
ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective Reasoning 1 1 13 Mar 2025 not harvested
SCOT: Self-Supervised Contrastive Pretraining For Zero-Shot Compositional Retrieval 0 1 12 Jan 2025 not harvested
MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval 1 1 19 Dec 2024 not harvested
Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval 1 3 15 Dec 2024 not harvested
Imagine and Seek: Improving Composed Image Retrieval with an Imagined Proxy 0 2 24 Nov 2024 not harvested
Semantic Editing Increment Benefits Zero-Shot Composed Image Retrieval 2 3 28 Oct 2024 not harvested
Training-free Zero-shot Composed Image Retrieval via Weighted Modality Fusion and Similarity 1 4 7 Sep 2024 ran 2 of 2 samples (0 unverified)
LDRE: LLM-based Divergent Reasoning and Ensemble for Zero-Shot Composed Image Retrieval 2 3 11 Jul 2024 not harvested
An Efficient Post-hoc Framework for Reducing Task Discrepancy of Text Encoders for Composed Image Retrieval 1 2 13 Jun 2024 not harvested
VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval 1 1 6 Jun 2024 ran 0 of 2 samples (2 unverified)
CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval 1 1 29 May 2024 not harvested
iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval 2 4 5 May 2024 ran 3 of 4 samples (1 unverified; 4 pointer-only for licence)
Improving Composed Image Retrieval via Contrastive Learning with Scaling Positives and Negatives 1 2 17 Apr 2024 ran 3 of 4 samples (1 unverified)
MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions 1 4 28 Mar 2024 ran 1 of 1 samples (0 unverified)
Language-only Efficient Training of Zero-shot Composed Image Retrieval 1 2 4 Dec 2023 ran 3 of 8 samples (5 unverified; 8 pointer-only for licence)
Pretrain like Your Inference: Masked Tuning Improves Zero-Shot Composed Image Retrieval 1 2 13 Nov 2023 ran 3 of 4 samples (1 unverified; 4 pointer-only for licence)
Vision-by-Language for Training-Free Compositional Image Retrieval 1 3 13 Oct 2023 ran 4 of 8 samples (4 unverified)
Sentence-level Prompts Benefit Composed Image Retrieval 1 2 9 Oct 2023 ran 7 of 8 samples (1 unverified; 8 pointer-only for licence)
Context-I2W: Mapping Images to Context-dependent Words for Accurate Zero-Shot Composed Image Retrieval 1 1 28 Sep 2023 ran 1 of 1 samples (0 unverified)
CoVR-2: Automatic Data Construction for Composed Video Retrieval 1 2 28 Aug 2023 ran 1 of 1 samples (0 unverified)
Composed Image Retrieval using Contrastive Learning and Task-oriented CLIP-based Features 2 1 22 Aug 2023 ran 4 of 5 samples (1 unverified)
Zero-shot Composed Text-Image Retrieval 1 1 12 Jun 2023 not harvested
Candidate Set Re-ranking for Composed Image Retrieval with Dual Multi-modal Encoder 2 1 25 May 2023 ran 4 of 6 samples (2 unverified)
Bi-directional Training for Composed Image Retrieval via Text Prompt Learning 1 1 29 Mar 2023 not harvested
Zero-Shot Composed Image Retrieval with Textual Inversion 2 2 27 Mar 2023 not harvested
CompoDiff: Versatile Composed Image Retrieval With Latent Diffusion 1 2 21 Mar 2023 ran 3 of 6 samples (3 unverified)
Data Roaming and Quality Assessment for Composed Image Retrieval 1 2 16 Mar 2023 not harvested
Pic2Word: Mapping Pictures to Words for Zero-shot Composed Image Retrieval 1 1 6 Feb 2023 ran 1 of 1 samples (0 unverified)

The full list of 35 is in the JSON twin.

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

MIT License

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • CIRR

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections