{"url":"/dataset/cirr","name":"CIRR","full_name":"Compose Image Retrieval on Real-life images","description_markdown":"**Composed Image Retrieval** (or, **Image Retreival conditioned on Language Feedback**) is a relatively new retrieval task, where an input query consists of an image and short textual description of how to modify the image. \r\n\r\nFor humans, the advantage of a bi-modal query is clear: some concepts and attributes are more succinctly described visually, others through language. By cross-referencing the two modalities, a reference image can capture the general gist of a scene, while the text can specify finer details.\r\n\r\nWe identify a major challenge of this task as the inherent ambiguity in knowing what information is important (typically one object of interest in the scene) and what can be ignored (e.g., the background and other irrelevant objects).\r\n\r\nWe release the first dataset of open-domain, real-life images with human-generated modification sentences, which support research on one-shot composed image retrieval, dialogue systems, fine-grained visiolinguistic reasoning, and more.","description_withheld":null,"homepage":"https://cuberick-orion.github.io/CIRR/","introduced_date":"2021-08-09","introduced_date_note":null,"introduced_by":{"paper":"/paper/image-retrieval-on-real-life-images-with-pre","title":"Image Retrieval on Real-life Images with Pre-trained Vision-and-Language Models","first_author":"Zheyuan Liu","url":null},"license":{"name":"MIT License","url":"https://github.com/Cuberick-Orion/CIRR/blob/main/LICENSE"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Image Retrieval","url":"/task/image-retrieval","datasets_with_task":"/datasets/task/image-retrieval"},{"name":"Composed Image Retrieval (CoIR)","url":"/task/composed-image-retrieval","datasets_with_task":"/datasets/task/composed-image-retrieval"},{"name":"Zero-Shot Composed Image Retrieval (ZS-CIR)","url":"/task/zero-shot-composed-image-retrieval-zs-cir","datasets_with_task":"/datasets/task/zero-shot-composed-image-retrieval-zs-cir"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["CIRR"],"data_loaders":[{"repo":"https://github.com/Cuberick-Orion/CIRR","url":"https://github.com/Cuberick-Orion/CIRR","frameworks":[]}],"num_papers_in_archive":61,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/zero-shot-composed-image-retrieval-zs-cir-on-1","task":"Zero-Shot Composed Image Retrieval (ZS-CIR)","dataset_variant":"CIRR","rows":47,"metrics":["R@1","R@5","R@10","R@50","Rsubset@1"],"first_row_in_archive_order":{"model":"CoLLM (finetuned - BLIP-L/16)","paper":"/paper/collm-a-large-language-model-for-composed","metrics":{"R@1":"45.8","R@10":"84.7","R@50":"95.8"},"code_links":[{"title":"hmchuong/CoLLM","url":"https://github.com/hmchuong/CoLLM"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/image-retrieval-on-cirr","task":"Image Retrieval","dataset_variant":"CIRR","rows":17,"metrics":["(Recall@5+Recall_subset@1)/2","Recall@10"],"first_row_in_archive_order":{"model":"TMCIR","paper":"/paper/tmcir-token-merge-benefits-composed-image","metrics":{"(Recall@5+Recall_subset@1)/2":"83.46","Recall@10":"91.06"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/composed-image-retrieval-coir-on-cirr-1","task":"Composed Image Retrieval (CoIR)","dataset_variant":"CIRR","rows":1,"metrics":["R@1","R@5"],"first_row_in_archive_order":{"model":"CoVR-BLIP-2","paper":"/paper/covr-learning-composed-video-retrieval-from","metrics":{"R@1":"50.43","R@5":"81.08"},"code_links":[{"title":"lucas-ventura/CoVR","url":"https://github.com/lucas-ventura/CoVR"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/tmcir-token-merge-benefits-composed-image","title":"TMCIR: Token Merge Benefits Composed Image Retrieval","date":"2025-04-15","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/collm-a-large-language-model-for-composed","title":"CoLLM: A Large Language Model for Composed Image Retrieval","date":"2025-03-25","rows_on_this_dataset":3,"code_links":1,"syntology":null},{"paper":"/paper/imagescope-unifying-language-guided-image-1","title":"ImageScope: Unifying Language-Guided Image Retrieval via Large Multimodal Model Collective Reasoning","date":"2025-03-13","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/scot-self-supervised-contrastive-pretraining","title":"SCOT: Self-Supervised Contrastive Pretraining For Zero-Shot Compositional Retrieval","date":"2025-01-12","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/megapairs-massive-data-synthesis-for","title":"MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval","date":"2024-12-19","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/reason-before-retrieve-one-stage-reflective","title":"Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval","date":"2024-12-15","rows_on_this_dataset":3,"code_links":1,"syntology":null},{"paper":"/paper/imagine-and-seek-improving-composed-image","title":"Imagine and Seek: Improving Composed Image Retrieval with an Imagined Proxy","date":"2024-11-24","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/semantic-editing-increment-benefits-zero-shot","title":"Semantic Editing Increment Benefits Zero-Shot Composed Image Retrieval","date":"2024-10-28","rows_on_this_dataset":3,"code_links":2,"syntology":null},{"paper":"/paper/training-free-zs-cir-via-weighted-modality","title":"Training-free Zero-shot Composed Image Retrieval via Weighted Modality Fusion and Similarity","date":"2024-09-07","rows_on_this_dataset":4,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":2,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/ldre-llm-based-divergent-reasoning-and","title":"LDRE: LLM-based Divergent Reasoning and Ensemble for Zero-Shot Composed Image Retrieval","date":"2024-07-11","rows_on_this_dataset":3,"code_links":2,"syntology":null},{"paper":"/paper/reducing-task-discrepancy-of-text-encoders","title":"An Efficient Post-hoc Framework for Reducing Task Discrepancy of Text Encoders for Composed Image Retrieval","date":"2024-06-13","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/vista-visualized-text-embedding-for-universal","title":"VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval","date":"2024-06-06","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":0,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/cala-complementary-association-learning-for","title":"CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval","date":"2024-05-29","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/isearle-improving-textual-inversion-for-zero","title":"iSEARLE: Improving Textual Inversion for Zero-Shot Composed Image Retrieval","date":"2024-05-05","rows_on_this_dataset":4,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":3,"samples_unverified":1,"pointer_only_for_licence":4,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/improving-composed-image-retrieval-via","title":"Improving Composed Image Retrieval via Contrastive Learning with Scaling Positives and Negatives","date":"2024-04-17","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":3,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/magiclens-self-supervised-image-retrieval","title":"MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions","date":"2024-03-28","rows_on_this_dataset":4,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/language-only-efficient-training-of-zero-shot","title":"Language-only Efficient Training of Zero-shot Composed Image Retrieval","date":"2023-12-04","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":8,"samples_ran":3,"samples_unverified":5,"pointer_only_for_licence":8,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/pretrain-like-you-inference-masked-tuning","title":"Pretrain like Your Inference: Masked Tuning Improves Zero-Shot Composed Image Retrieval","date":"2023-11-13","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":3,"samples_unverified":1,"pointer_only_for_licence":4,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/vision-by-language-for-training-free","title":"Vision-by-Language for Training-Free Compositional Image Retrieval","date":"2023-10-13","rows_on_this_dataset":3,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":8,"samples_ran":4,"samples_unverified":4,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/sentence-level-prompts-benefit-composed-image","title":"Sentence-level Prompts Benefit Composed Image Retrieval","date":"2023-10-09","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":8,"samples_ran":7,"samples_unverified":1,"pointer_only_for_licence":8,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/context-i2w-mapping-images-to-context","title":"Context-I2W: Mapping Images to Context-dependent Words for Accurate Zero-Shot Composed Image Retrieval","date":"2023-09-28","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/covr-learning-composed-video-retrieval-from","title":"CoVR-2: Automatic Data Construction for Composed Video Retrieval","date":"2023-08-28","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/composed-image-retrieval-using-contrastive","title":"Composed Image Retrieval using Contrastive Learning and Task-oriented CLIP-based Features","date":"2023-08-22","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":5,"samples_ran":4,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/zero-shot-composed-text-image-retrieval","title":"Zero-shot Composed Text-Image Retrieval","date":"2023-06-12","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/candidate-set-re-ranking-for-composed-image","title":"Candidate Set Re-ranking for Composed Image Retrieval with Dual Multi-modal Encoder","date":"2023-05-25","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":4,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/bi-directional-training-for-composed-image","title":"Bi-directional Training for Composed Image Retrieval via Text Prompt Learning","date":"2023-03-29","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/zero-shot-composed-image-retrieval-with","title":"Zero-Shot Composed Image Retrieval with Textual Inversion","date":"2023-03-27","rows_on_this_dataset":2,"code_links":2,"syntology":null},{"paper":"/paper/compodiff-versatile-composed-image-retrieval","title":"CompoDiff: Versatile Composed Image Retrieval With Latent Diffusion","date":"2023-03-21","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":3,"samples_unverified":3,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/data-roaming-and-early-fusion-for-composed","title":"Data Roaming and Quality Assessment for Composed Image Retrieval","date":"2023-03-16","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/pic2word-mapping-pictures-to-words-for-zero","title":"Pic2Word: Mapping Pictures to Words for Zero-shot Composed Image Retrieval","date":"2023-02-06","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/conditioned-and-composed-image-retrieval","title":"Conditioned and Composed Image Retrieval Combining and Partially Fine-Tuning CLIP-Based Features","date":"2022-06-19","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/this-is-my-unicorn-fluffy-personalizing","title":"\"This is my unicorn, Fluffy\": Personalizing frozen vision-language representations","date":"2022-04-04","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/artemis-attention-based-retrieval-with-text-1","title":"ARTEMIS: Attention-based Retrieval with Text-Explicit Matching and Implicit Similarity","date":"2022-03-15","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/effective-conditioned-and-composed-image","title":"Effective Conditioned and Composed Image Retrieval Combining CLIP-Based Features","date":"2022-01-01","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/image-retrieval-on-real-life-images-with-pre","title":"Image Retrieval on Real-life Images with Pre-trained Vision-and-Language Models","date":"2021-08-09","rows_on_this_dataset":1,"code_links":3,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":13,"samples_ran":0,"samples_unverified":13,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":17,"samples_harvested":75,"samples_ran":41,"samples_unverified":34,"pointer_only_for_licence":25,"papers_with_no_sample_that_ran":2,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}