{"url":"/dataset/crepe-vision-language","name":"CREPE (Compositional REPresentation Evaluation)","full_name":null,"description_markdown":"A fundamental characteristic common to both human vision and natural language is their compositional nature. Yet,\r\ndespite the performance gains contributed by large vision\r\nand language pretraining, we find that—across 7 architectures trained with 4 algorithms on massive datasets—they\r\nstruggle at compositionality. To arrive at this conclusion, we\r\nintroduce a new compositionality evaluation benchmark,\r\nCREPE, which measures two important aspects of compositionality identified by cognitive science literature: systematicity and productivity. To measure systematicity, CREPE\r\nconsists of a test dataset containing over 370K image-text\r\npairs and three different seen-unseen splits. The three splits\r\nare designed to test models trained on three popular training\r\ndatasets: CC-12M, YFCC-15M, and LAION-400M. We also\r\ngenerate 325K, 316K, and 309K hard negative captions\r\nfor a subset of the pairs. To test productivity, CREPE contains 17K image-text pairs with nine different complexities\r\nplus 183K hard negative captions with atomic, swapping\r\nand negation foils. The datasets are generated by repurposing the Visual Genome scene graphs and region descriptions\r\nand applying handcrafted templates and GPT-3. For systematicity, we find that model performance decreases consistently when novel compositions dominate the retrieval set,\r\nwith Recall@1 dropping by up to 12%. For productivity,\r\nmodels’ retrieval success decays as complexity increases,\r\nfrequently nearing random chance at high complexity. These\r\nresults hold regardless of model and training dataset size.","description_withheld":null,"homepage":"","introduced_date":"2022-12-13","introduced_date_note":null,"introduced_by":{"paper":"/paper/crepe-can-vision-language-foundation-models","title":"CREPE: Can Vision-Language Foundation Models Reason Compositionally?","first_author":"Zixian Ma","url":null},"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Image Retrieval","url":"/task/image-retrieval","datasets_with_task":"/datasets/task/image-retrieval"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["CREPE (Compositional REPresentation Evaluation)"],"data_loaders":[],"num_papers_in_archive":2,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/image-retrieval-on-crepe-vision-language","task":"Image Retrieval","dataset_variant":"CREPE (Compositional REPresentation Evaluation)","rows":22,"metrics":["Recall@1 (HN-Atom + HN-Comp, SC)","Recall@1 (HN-Atom + HN-Comp, UC)","Recall@1 (HN-Atom, UC)","Recall@1 (HN-Comp, UC)"],"first_row_in_archive_order":{"model":"ViT-L-14 (LAION400M)","paper":"/paper/crepe-can-vision-language-foundation-models","metrics":{"Recall@1 (HN-Atom + HN-Comp, SC)":"39.44","Recall@1 (HN-Atom + HN-Comp, UC)":"33.81","Recall@1 (HN-Atom, UC)":"47.86","Recall@1 (HN-Comp, UC)":"60.78"},"code_links":[{"title":"raivnlab/crepe","url":"https://github.com/raivnlab/crepe"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/coarse-to-fine-contrastive-learning-in-image","title":"Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language Compositionality","date":"2023-05-23","rows_on_this_dataset":14,"code_links":0,"syntology":null},{"paper":"/paper/crepe-can-vision-language-foundation-models","title":"CREPE: Can Vision-Language Foundation Models Reason Compositionally?","date":"2022-12-13","rows_on_this_dataset":8,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":3,"samples_unverified":0,"pointer_only_for_licence":3,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":1,"samples_harvested":3,"samples_ran":3,"samples_unverified":0,"pointer_only_for_licence":3,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}