{"url":"/dataset/inpaintcoco","name":"InpaintCOCO","full_name":null,"description_markdown":"InpaintCOCO is a benchmark to understand fine-grained concepts in multimodal models (vision-language) similar to [Winoground](https://arxiv.org/abs/2204.03162). To our knowledge InpaintCOCO is the first benchmark, which consists of image pairs with minimum differences, so that the visual representation can be analyzed in a more standardized setting.\r\n\r\nA data sample contains 2 images and 2 corresponding captions that differ only in one object, the color of an object, or the size of an object. \r\n\r\nThe metric used in the paper compares if true image-text pairs are more similar than wrong image-text combinations and that for both image-text pairs:\r\n\\begin{equation}\r\n\\begin{split}\r\n\\operatorname{sim}(i_\\text{COCO},t_\\text{COCO}) > \\operatorname{sim}(i_\\text{inp},t_\\text{COCO}) \\quad \\land \\\\\r\n\\operatorname{sim}(i_\\text{inp},t_\\text{inp}) > \\operatorname{sim}(i_\\text{COCO},t_\\text{inp})\r\n\\end{split}\r\n\\label{eq:challengeset}\r\n\\end{equation}\r\n\r\nInpaintCOCO is published in [Proceedings of the 3rd Workshop on Advances in Language and Vision Research (ALVR)](https://aclanthology.org/2024.alvr-1.9/) at [ACL 2024](https://aclanthology.org/volumes/2024.alvr-1/).","description_withheld":null,"homepage":"https://huggingface.co/datasets/phiyodr/InpaintCOCO","introduced_date":"2024-08-16","introduced_date_note":null,"introduced_by":null,"license":{"name":"diverse licenses","url":"https://huggingface.co/datasets/phiyodr/InpaintCOCO#licensing-information"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Fine-Grained Image Inpainting","url":"/task/fine-grained-image-inpainting","datasets_with_task":"/datasets/task/fine-grained-image-inpainting"},{"name":"Image-text Retrieval","url":"/task/image-text-retrieval","datasets_with_task":"/datasets/task/image-text-retrieval"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["InpaintCOCO"],"data_loaders":[],"num_papers_in_archive":0,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}