Datasets › CIRR
CIRR (Compose Image Retrieval on Real-life images)
Composed Image Retrieval (or, Image Retreival conditioned on Language Feedback) is a relatively new retrieval task, where an input query consists of an image and short textual description of how to modify the image.
For humans, the advantage of a bi-modal query is clear: some concepts and attributes are more succinctly described visually, others through language. By cross-referencing the two modalities, a reference image can capture the general gist of a scene, while the text can specify finer details.
We identify a major challenge of this task as the inherent ambiguity in knowing what information is important (typically one object of interest in the scene) and what can be ignored (e.g., the background and other irrelevant objects).
We release the first dataset of open-domain, real-life images with human-generated modification sentences, which support research on one-shot composed image retrieval, dialogue systems, fine-grained visiolinguistic reasoning, and more.
Benchmarks archive 2025-07-28
All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Zero-Shot Composed Image Retrieval (ZS-CIR) | CIRR | CoLLM (finetuned - BLIP-L/16) R@1 45.8 | CoLLM: A Large Language Model for Composed Image Retrieval | hmchuong/CoLLM | 47 | Compare |
| Image Retrieval | CIRR | TMCIR (Recall@5+Recall_subset@1)/2 83.46 | TMCIR: Token Merge Benefits Composed Image Retrieval | — | 17 | Compare |
| Composed Image Retrieval (CoIR) | CIRR | CoVR-BLIP-2 R@1 50.43 | CoVR-2: Automatic Data Construction for Composed Video Retrieval | lucas-ventura/CoVR | 1 | Compare |
Papers archive 2025-07-28
30 shown of 35 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 61. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
The full list of 35 is in the JSON twin.
Dataset loaders archive 2025-07-28
1 loader as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- CIRR
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections